Managed multi-model AI inference platform
Built and hardened a managed, multi model AI inference platform with an OpenAI compatible API, provider routing, verified model catalogs, and centralized operational controls.
The platform brings together Clerk authenticated workspaces, tiered Stripe subscriptions, rolling SGD credit accounting, usage tracking, API key management, and administrator controls. It also supports Telegram bots with one configured model per bot, isolated conversation history, reset controls, image input, voice transcription, streaming responses, managed webhooks, and bounded reply delivery.
Web search is available through user owned Exa or Brave credentials, with encrypted storage, provider verification, exact URL retrieval, hostname restricted fallback search, failure safe handling, and only one active provider per user. The application also includes English, Simplified Chinese, and Malay localization.
Security and reliability controls cover shared inference quotas, concurrent provider work, expiring database leases, browser access boundaries, pinned HTTPS transport, normalized image handling, account suspension, subscription grant idempotency, and safe migration replay.
The result is a practical platform for serving multiple AI models through one consistent API while maintaining control over access, billing, usage, security, and administration. Internally, the service runs through a controlled local compute environment based in Singapore for operational testing and validation.