All work
Agents
Multi-agent LLM Gateway
Agents that share one gateway to Bedrock, OpenAI and Gemini — with caching, cost tracking and MCP tools.
The problem
When every agent talks to LLM providers on its own, API keys spread everywhere, costs are invisible and switching models means rewriting code.
How it works
A central LLM gateway puts three providers behind one API, with a TTL cache and per-request cost, token and latency metrics. LangGraph agents call tools through an MCP server and are exposed over REST and WebSocket streaming. Every service runs in Docker, ready for Kubernetes.
Pipeline
- 01Client
- 02Agent (REST / WS)
- 03MCP tools
- 04LLM gateway
- 05Bedrock · OpenAI · Gemini
Highlights
- Switch providers without touching agent code
- Response cache that cuts repeated API spend
- Real-time streaming over WebSocket
- Health checks on every service; ready for AWS EKS
Stack
LangGraphMCPFastAPIAWS BedrockOpenAIGeminiDockerKubernetes
Want something like this for your company?
I build production versions of these systems on your data and your infrastructure.
Related service: Agents & workflow automation
Book a call