Comprehensive architecture documentation for Kolya BR Proxy -- an AI Gateway that provides both OpenAI-compatible and Anthropic Messages API access to AWS Bedrock models (Claude, Nova, DeepSeek, Mistral, Llama, etc.), Google Gemini models via the native generateContent API, and OpenAI GPT-5.5 / GPT-5.4 models served by the AWS "mantle" inference engine via the OpenAI Responses API.
- System Architecture Overview
- Backend Layer Architecture
- Database ER Diagram
- Frontend Architecture
- Infrastructure Architecture
- Authentication Flow
- Request Processing Flow
- Pricing Model
The system follows a classic three-tier architecture: a Vue 3 frontend served by Nginx, a FastAPI backend running on Uvicorn, and AWS Bedrock as the upstream LLM provider. All components run inside an AWS EKS cluster with PostgreSQL for persistence, Redis for distributed rate limiting, and External Secrets Operator (ESO) for secrets management.
graph LR
subgraph Clients ["Client Applications"]
OAI["OpenAI SDK / Client Libraries"]
Anthropic_SDK["Anthropic SDK / Client Libraries"]
Browser["Admin Dashboard<br/>(Browser)"]
end
subgraph AWS_EKS ["AWS EKS Cluster"]
subgraph Frontend_Pod ["Frontend Pod"]
Nginx["Nginx"]
Vue["Vue 3 / Quasar SPA"]
end
subgraph Backend_Pod ["Backend Pod"]
Uvicorn["Uvicorn ASGI Server"]
FastAPI["FastAPI Application"]
end
ALB["Application Load<br/>Balancer (ALB)"]
end
subgraph Data_Stores ["Data Stores"]
PG["Aurora PostgreSQL"]
Redis["Redis Standalone<br/>(distributed rate limiting)"]
end
subgraph AWS_Services ["AWS Services"]
Bedrock["AWS Bedrock<br/>(Claude, Nova, DeepSeek,<br/>Mistral, Llama)"]
end
subgraph Google_Services ["Google Services"]
Gemini["Google Gemini API<br/>(generateContent native)"]
end
subgraph Mantle_Services ["AWS mantle (OpenAI models)"]
Mantle["mantle Inference Engine<br/>(OpenAI Responses API)<br/>(GPT-5.5, GPT-5.4)"]
end
OAI -->|"HTTPS /v1/chat/*"| ALB
Anthropic_SDK -->|"HTTPS /v1/messages"| ALB
Browser -->|"HTTPS /*"| ALB
ALB -->|"Frontend routes"| Nginx
ALB -->|"API routes"| Uvicorn
Nginx --> Vue
Uvicorn --> FastAPI
FastAPI -->|"InvokeModel /<br/>Converse API"| Bedrock
FastAPI -->|"generateContent /<br/>streamGenerateContent"| Gemini
FastAPI -->|"SigV4 POST /responses<br/>(OpenAI Responses API)"| Mantle
FastAPI -->|"SQLAlchemy async"| PG
FastAPI -.->|"Distributed rate limiting"| Redis
| Decision | Rationale |
|---|---|
| Dual API compatibility | OpenAI-compatible (/v1/chat/completions) and Anthropic Messages API (/v1/messages); clients only change base_url and api_key |
| Configurable API key prefix | Keys default to kbr_ prefix; sk-ant-api03 prefix available for Claude Code / Anthropic SDK compatibility |
| Strip thinking blocks from history | Bedrock doesn't support adaptive signature-only thinking blocks, so they are removed from conversation history before forwarding |
| Dynamic model resolution | _ProfileCache queries AWS APIs at startup + daily 03:00 UTC to discover available inference profiles and foundation models; resolve_model() routes dynamically instead of using hardcoded prefix lists |
Singleton BedrockClient |
One shared aioboto3 session + connection pool per process |
| Asyncio semaphore (50) | Back-pressure to match connection pool size; prevents request queuing |
| JWT for dashboard, API keys for gateway | Separate auth concerns; API keys are long-lived, JWTs are short-lived |
| Background usage recording | record_usage runs as a background task to avoid blocking responses |
| Gemini native API (not OpenAI compat) | Uses generateContent / streamGenerateContent directly; avoids Gemini's OpenAI-compat layer which rejects fields like frequency_penalty |
| Gemini format conversion in client layer | GeminiClient converts OpenAI ↔ Gemini natively; chat.py sees uniform OpenAI format regardless of backend |
| mantle reuses the Gemini provider pattern | OpenAI GPT-5.5/5.4 (served by AWS mantle) are intercepted by early routing + a dedicated MantleClient + pricing short-circuit before reaching BedrockClient, exactly like Gemini; this leaves the existing Claude/Nova/Gemini paths untouched. is_openai_mantle_model() does an exact match on openai.gpt-5.5 / openai.gpt-5.4 (avoiding false hits on open-source gpt-oss) |
| mantle uses the OpenAI Responses API over SigV4 | mantle is reached at https://bedrock-mantle.{region}.api.aws/openai/v1 via POST /responses (not boto3 converse/invoke_model). Requests are SigV4-signed (botocore SigV4Auth/AWSRequest, service name bedrock) reusing the existing AWS credential chain (EKS Pod IRSA, incl. SessionToken); MantleClient converts OpenAI ChatCompletions ↔ Responses natively |
The backend is organized into four layers: API (routing + validation), Middleware (security + CORS), Service (business logic), and Data (SQLAlchemy models). The entry point is backend/main.py which sets up the FastAPI app via the create_app() factory and the lifespan context manager.
graph TD
subgraph API_Layer ["API Layer (Routers)"]
Health["/health<br/>health_router"]
Admin["/admin<br/>admin_router"]
V1["/v1<br/>gateway_router"]
subgraph Admin_Sub ["Admin Endpoints"]
Auth_EP["/admin/auth<br/>OAuth (Cognito, Microsoft), refresh"]
Tokens_EP["/admin/tokens<br/>CRUD API keys"]
Usage_EP["/admin/usage<br/>usage statistics"]
Audit_EP["/admin/audit-logs<br/>audit log queries"]
Models_EP["/admin/models<br/>model management"]
end
subgraph V1_Sub ["Gateway Endpoints"]
subgraph OpenAI_Sub ["OpenAI-compatible"]
Chat_EP["/v1/chat/completions<br/>streaming + non-streaming"]
Models_List["/v1/models<br/>list available models"]
end
subgraph Anthropic_Sub ["Anthropic Messages API"]
Messages_EP["/v1/messages<br/>streaming + non-streaming"]
end
end
Admin --> Admin_Sub
V1 --> V1_Sub
end
subgraph Middleware_Layer ["Middleware Layer"]
direction LR
Cache_MW["Cache-Control<br/>Middleware"]
Security_MW["SecurityMiddleware<br/>(CSRF, origin, headers)"]
CORS_MW["CORSMiddleware<br/>(FastAPI built-in)"]
end
subgraph Service_Layer ["Service Layer"]
BedrockSvc["BedrockClient<br/>(singleton, semaphore=50,<br/>_ProfileCache)"]
ReqTranslator["RequestTranslator<br/>(OpenAI -> Bedrock)"]
ResTranslator["ResponseTranslator<br/>(Bedrock -> OpenAI)"]
AnthReqTranslator["AnthropicRequestTranslator<br/>(Anthropic -> Bedrock)"]
AnthResTranslator["AnthropicResponseTranslator<br/>(Bedrock -> Anthropic)"]
GeminiSvc["GeminiClient<br/>(OpenAI ↔ Gemini native conversion)"]
GeminiPricing["GeminiPricingUpdater<br/>(3-tier price fetch)"]
MantleSvc["MantleClient<br/>(SigV4; OpenAI ↔ Responses conversion;<br/>native /responses passthrough)"]
TokenSvc["TokenService<br/>(validate, CRUD)"]
AuthSvc["AuthService<br/>(OAuth user creation)"]
RefreshSvc["RefreshTokenService<br/>(token rotation)"]
AuditSvc["AuditLogService<br/>(security events)"]
PricingSvc["ModelPricing<br/>(cost calculation)"]
BGTasks["BackgroundTaskManager<br/>(usage recording)"]
OAuthSvc["MicrosoftOAuth /<br/>CognitoOAuth"]
end
subgraph Data_Layer ["Data Layer (SQLAlchemy Async Models)"]
UserModel["User"]
APITokenModel["APIToken"]
ModelModel["Model"]
TeamModel["Team"]
TeamMemberModel["TeamMember"]
UsageModel["UsageRecord"]
AuditModel["AuditLog"]
RefreshModel["RefreshToken"]
OAuthModel["OAuthState"]
SysConfig["SystemConfig"]
end
API_Layer --> Middleware_Layer
Middleware_Layer --> Service_Layer
Service_Layer --> Data_Layer
Middleware is registered in create_app() in backend/main.py. FastAPI processes middleware in reverse registration order (last added = outermost). The effective processing order for an incoming request is:
- Cache-Control -- adds
no-cache, no-storeheaders to every response - SecurityMiddleware -- origin validation, CSRF protection (
X-Requested-With), security response headers (X-Content-Type-Options,X-Frame-Options, CSP) - CORSMiddleware -- handles
OPTIONSpreflight, setsAccess-Control-*headers
Every log line is automatically tagged with the API key name using Python contextvars. When a gateway request is authenticated, the token_name is set in a context variable and injected into all subsequent log output for that request.
- Log format:
%(asctime)s - %(name)s - %(levelname)s - [%(token_name)s] %(message)s - Filtering:
grep '\[key-name\]'isolates all activity for a specific API key across all log files - Scope: Applies to all gateway request processing, including service layer and background tasks
Streaming requests use a two-level failover mechanism to improve reliability when an upstream region or model is temporarily unavailable:
| Level | Strategy | Client Impact |
|---|---|---|
| Level 1 -- Cross-region | Retry the same model via a different AWS region (using cross-region inference profiles) | Transparent; client sees no difference |
| Level 2 -- Model degradation | Fall back to the next model in the configured fallback chain | Client notified via x-actual-model SSE comment in the stream |
How it works: After sending the request upstream, the service layer waits for the first content chunk. If no content arrives within the timeout, the current attempt is abandoned and the next level is tried. All SSE events buffered during the failed attempt are discarded, preventing metadata leakage from partial responses.
| Config | Default | Description |
|---|---|---|
STREAM_FIRST_CONTENT_TIMEOUT |
600s | Maximum wait time for the first content chunk before triggering failover |
STREAM_MODEL_FALLBACK_CHAIN |
(none) | Ordered list of fallback models for Level 2 degradation |
# backend/main.py
@asynccontextmanager
async def lifespan(app: FastAPI):
await init_db() # Initialize database connection pool
# Initialize pricing data if DB is empty (AWS + Gemini)
bedrock = BedrockClient.get_instance() # Create singleton Bedrock client
await bedrock.refresh_profile_cache() # Populate inference profile cache from AWS APIs
start_scheduler() # APScheduler: pricing @ 02:00/02:30, profile cache @ 03:00 UTC
yield
stop_scheduler()
# Shutdown: cleanup resourcesAll models use UUID primary keys and are defined in backend/app/models/. Relationships are enforced via SQLAlchemy ORM with cascade deletes where appropriate.
erDiagram
User ||--o{ APIToken : "owns"
User ||--o{ UsageRecord : "generates"
User ||--o{ RefreshToken : "has"
User ||--o{ AuditLog : "triggers"
User ||--o{ Team : "manages"
APIToken ||--o{ UsageRecord : "tracks"
APIToken ||--o{ Model : "enables"
APIToken ||--o| TeamMember : "belongs to"
Team ||--o{ TeamMember : "has"
RefreshToken ||--o{ RefreshToken : "parent-child"
User {
uuid id PK
string email UK
string password_hash "nullable (unused, OAuth-only)"
enum auth_method "LOCAL | MICROSOFT | COGNITO"
boolean is_active
boolean is_admin
enum role "super_admin | admin"
json permissions "RBAC permission scopes"
boolean email_verified
decimal current_balance "Numeric(10,2)"
string microsoft_id UK "nullable"
string first_name
string last_name
datetime created_at
datetime updated_at
datetime last_login_at
}
APIToken {
uuid id PK
uuid user_id FK
string name
string description "nullable"
string token_hash UK "SHA256"
string encrypted_token "Fernet AES-128"
datetime expires_at "nullable"
decimal quota_usd "Numeric(10,2) nullable"
decimal monthly_quota_usd "Numeric(10,2) nullable"
string monthly_reset_policy "nullable (reset | rollover)"
datetime monthly_quota_start "nullable"
string_array allowed_ips "nullable"
boolean is_active
boolean is_deleted
json token_metadata
datetime created_at
datetime updated_at
datetime last_used_at
datetime deleted_at
}
Team {
uuid id PK
uuid user_id FK
string name
decimal monthly_budget_usd "Numeric(10,2)"
string monthly_reset_policy "reset | rollover"
boolean daily_limit_enabled
datetime monthly_budget_start
boolean is_active
datetime created_at
datetime updated_at
}
TeamMember {
uuid id PK
uuid team_id FK "CASCADE"
uuid token_id FK "CASCADE, unique"
decimal allocated_usd "Numeric(10,2)"
datetime created_at
datetime updated_at
}
Model {
uuid id PK
uuid token_id FK "CASCADE"
string model_name
boolean is_active
boolean is_deleted
datetime created_at
datetime updated_at
datetime deleted_at
}
UsageRecord {
uuid id PK
uuid user_id FK
uuid token_id FK
string request_id
string model
int prompt_tokens
int completion_tokens
int total_tokens
int cache_creation_input_tokens "Cache write tokens"
int cache_read_input_tokens "Cache read tokens"
decimal cost_usd "Numeric(10,4)"
json request_metadata
datetime created_at
}
AuditLog {
uuid id PK
uuid user_id FK "nullable, SET NULL"
enum action "LOGIN_SUCCESS, TOKEN_CREATED, etc."
boolean success
text details "JSON string"
string error_message
string ip_address
string user_agent
string resource_type
string resource_id
datetime created_at
}
RefreshToken {
uuid id PK
uuid user_id FK "CASCADE"
string token_hash UK "SHA256"
uuid family_id "for theft detection"
uuid parent_token_id FK "self-ref, SET NULL"
datetime created_at
datetime expires_at
boolean is_revoked
datetime revoked_at
string revoked_reason
string ip_address
string user_agent
}
OAuthState {
uuid id PK
string state UK
string provider "microsoft | cognito"
string code_verifier "PKCE code_verifier (nullable)"
datetime created_at
datetime expires_at "10 min TTL"
}
SystemConfig {
uuid id PK
string key UK
text value
text description
boolean is_public
datetime created_at
datetime updated_at
}
- User.auth_method: Enum with values
LOCAL,MICROSOFT,COGNITO. OAuth users havepassword_hash = NULL. Therolefield (super_adminoradmin) andpermissionsJSON field implement RBAC. - APIToken: Stores both
token_hash(SHA256, for lookup) andencrypted_token(Fernet AES, for recovery). Thequota_usdfield limits total spending per token;monthly_quota_usdlimits monthly spending with configurablemonthly_reset_policy. Model access is controlled via the relatedModeltable rather than an array column. - Team / TeamMember: Teams group tokens under a shared monthly budget. Each
TeamMemberlinks a token to a team with anallocated_usdbudget. The token'smonthly_quota_usdis overridden by the team allocation when the token is a team member. - Model: Each row links one Bedrock model name to one APIToken. A token can access only models with
is_active=Trueandis_deleted=False. - RefreshToken.family_id: Groups related tokens for theft detection. If a revoked token is reused, the entire family is revoked.
The frontend is a Vue 3 SPA built with the Quasar framework (dark theme). It uses Pinia stores for state management, Vue Router for navigation with authentication guards, and Axios with automatic 401 refresh interceptors.
graph TD
subgraph Pages ["Pages (src/pages/)"]
Login["LoginPage"]
CognitoCallback["CognitoCallbackPage"]
MSCallback["MicrosoftCallbackPage"]
Dashboard["DashboardPage"]
Teams["TeamsPage"]
Tokens["TokensPage"]
Models["ModelsPage"]
Playground["PlaygroundPage"]
Monitor["MonitorPage"]
Activity["ActivityPage"]
AdminUsers["AdminUsersPage"]
Settings["SettingsPage"]
NotFound["ErrorNotFound"]
end
subgraph Layout ["Layout"]
MainLayout["MainLayout.vue<br/>(sidebar nav, user menu,<br/>dark theme shell)"]
end
subgraph Stores ["Pinia Stores (src/stores/)"]
AuthStore["auth.ts<br/>OAuth redirect, logout, JWT mgmt"]
TokenStore["tokens.ts<br/>API key CRUD"]
ModelStore["models.ts<br/>model management"]
DashStore["dashboard.ts<br/>usage overview stats"]
MonitorStore["monitor.ts<br/>usage charts & analytics"]
end
subgraph Boot ["Boot (src/boot/)"]
Axios["axios.ts<br/>API client, 401 interceptor,<br/>auto token refresh"]
end
subgraph Router ["Router"]
Routes["routes.ts<br/>requiresAuth guards"]
end
Router -->|"requiresAuth: true"| MainLayout
Router -->|"requiresAuth: false"| Login
Router -->|"requiresAuth: false"| CognitoCallback
Router -->|"requiresAuth: false"| MSCallback
MainLayout --> Dashboard
MainLayout --> Teams
MainLayout --> Tokens
MainLayout --> Models
MainLayout --> Playground
MainLayout --> Monitor
MainLayout --> Activity
MainLayout --> AdminUsers
MainLayout --> Settings
Pages --> Stores
Stores --> Boot
Boot -->|"HTTP requests to /admin/*"| Backend["Backend API"]
| Path | Page | Auth Required | Description |
|---|---|---|---|
/login |
LoginPage | No | User login (OAuth provider selection) |
/auth/cognito/callback |
CognitoCallbackPage | No | Cognito OAuth callback |
/auth/microsoft/callback |
MicrosoftCallbackPage | No | Microsoft OAuth callback |
/ |
DashboardPage | Yes | Overview, usage stats |
/teams |
TeamsPage | Yes | Team budget management |
/tokens |
TokensPage | Yes | API key management |
/models |
ModelsPage | Yes | Model configuration |
/playground |
PlaygroundPage | Yes | Test conversations |
/monitor |
MonitorPage | Yes | Usage charts & analytics |
/activity |
ActivityPage | Yes | Activity feed (management operations) |
/admin-users |
AdminUsersPage | Yes (super_admin) | Admin user management |
/settings |
SettingsPage | Yes | Account settings |
The MainLayout.vue renders a persistent left drawer with these menu items: Dashboard, Teams, API Keys, Models, Playground, Monitor, Activity, Admin Users (super_admin only), Settings. The header shows the app title and a user menu (email, balance, settings, logout).
Infrastructure is defined in Terraform (iac/). All configuration is centralized in iac/terraform.tfvars as the single source of truth (account, region, domains, feature toggles). The deploy-all.sh script orchestrates the full deployment in 6 steps (0-5), while destroy.sh handles safe teardown. It provisions a VPC, EKS cluster with Karpenter autoscaling, Aurora PostgreSQL, Redis for distributed rate limiting, and optional WAF / Global Accelerator. Secrets are managed via AWS Secrets Manager with External Secrets Operator (ESO) syncing them into Kubernetes.
graph TD
subgraph AWS_Region ["AWS Region"]
subgraph VPC ["VPC Module"]
subgraph Public_Subnets ["Public Subnets"]
ALB["Application Load<br/>Balancer (ALB)<br/>(dynamically created)"]
NAT["NAT Gateway"]
end
subgraph Private_Subnets ["Private Subnets"]
subgraph EKS ["EKS Cluster (eks_karpenter module)"]
Karpenter["Karpenter<br/>Node Autoscaler"]
LBC["AWS Load Balancer<br/>Controller"]
FrontendPod["Frontend Pod<br/>(Vue3/Quasar + Nginx)"]
BackendPod["Backend Pod<br/>(FastAPI + Uvicorn)"]
end
Aurora["Aurora PostgreSQL<br/>(RDS Module)<br/>(private, no external access)"]
Redis_Pod["Redis Standalone<br/>(kbp namespace)"]
ESO["External Secrets<br/>Operator (ESO)"]
end
end
ECR["ECR<br/>(Container Registry)"]
Bedrock["AWS Bedrock"]
SecretsManager["AWS Secrets Manager"]
GA["Global Accelerator<br/>(optional)"]
end
Internet["Internet"] -->|"HTTPS"| GA
GA --> ALB
Internet -->|"HTTPS (direct)"| ALB
ALB --> FrontendPod
ALB --> BackendPod
BackendPod -->|"Via NAT"| Bedrock
BackendPod --> Aurora
BackendPod -.->|"Rate limiting"| Redis_Pod
ESO -->|"Sync secrets<br/>(refreshInterval: 1h)"| SecretsManager
Karpenter --> Private_Subnets
LBC -.->|"Creates & manages"| ALB
ECR -.->|"Pull images"| EKS
Secrets are stored in AWS Secrets Manager and automatically synced to Kubernetes Secrets by the External Secrets Operator (ESO). This replaces local secrets.yaml files and ensures secrets never exist in version control.
| Component | Role |
|---|---|
| AWS Secrets Manager | Single source of truth for all secrets (DB credentials, JWT keys, OAuth client secrets, etc.) |
| External Secrets Operator (ESO) | Runs in-cluster, watches ExternalSecret CRDs, and syncs secrets from AWS Secrets Manager to K8s Secrets |
| Pod Identity | ESO authenticates to AWS Secrets Manager via EKS Pod Identity (no static AWS credentials) |
deploy-all.sh Step 4 |
Pushes secrets to AWS Secrets Manager via aws secretsmanager put-secret-value (preserves existing values) |
Sync behavior:
refreshInterval: 1h-- ESO re-fetches secrets from Secrets Manager every hour- Secrets are created as standard Kubernetes Secrets, consumed by Pods via
envFromorenvreferences - Secret rotation in AWS Secrets Manager is automatically picked up within the refresh interval
A Redis standalone instance runs in the kbp (kolya-br-proxy) namespace to provide distributed rate limiting across all backend Pods.
| Aspect | Details |
|---|---|
| Deployment | Redis standalone in kbp namespace (Kubernetes Deployment + Service) |
| Purpose | Global token bucket rate limiting via atomic Lua scripts |
| Fallback | If Redis is unavailable, each Pod falls back to a local in-memory LocalTokenBucket (per-Pod rate limiting, not skip) |
| Access | Backend Pods connect via Kubernetes Service DNS (redis.kbp.svc.cluster.local) |
All deployment configuration is centralized in iac/terraform.tfvars. Both deploy-all.sh and destroy.sh read from and write to this file.
| Key | Description |
|---|---|
account / region |
AWS account ID and target region |
frontend_domain / api_domain |
Domain names (e.g. kbp.kolya.fun, api.kbp.kolya.fun) |
project_name / project_name_alias |
Resource naming (some resources use full name, others use alias) |
enable_waf |
WAF toggle (auto-enabled after ALBs are ready in Step 4) |
enable_global_accelerator |
Global Accelerator toggle (Step 5) |
enable_cognito |
Authentication provider toggle (Step 0 selection) |
cognito_allowed_email_domains |
Email domain whitelist for Cognito |
| Step | Command | What It Does |
|---|---|---|
| 0 | --step 0 |
Auto-detect account/region, select auth provider, configure domains → write terraform.tfvars |
| 1 | --step 1 |
terraform init + plan + apply (VPC, EKS, RDS, Cognito, etc.) |
| 2 | --step 2 |
Deploy Helm charts (ALB Controller, Karpenter, Metrics Server) |
| 3 | --step 3 |
Build Docker images, push to ECR (domains read from tfvars) |
| 4 | --step 4 |
Deploy K8s app (generate configs from tfvars, push secrets to SM, auto-enable WAF) |
| 5 | --step 5 |
Toggle Global Accelerator on/off |
- Verify AWS identity and confirm target (account, region, workspace)
- Initialize Terraform and select workspace
- Disable WAF/GA via
terraform apply(theirdata "aws_lb"lookups require ALBs to exist) - Clean up K8s resources (Ingress first → triggers ALB deletion, then ExternalSecrets, namespace)
terraform destroyto remove all remaining infrastructure
Important: K8s resources (especially Ingress/ALB) must be deleted before
terraform destroy, otherwise ALBs and target groups will block Terraform.
| Module | Source Path | Purpose |
|---|---|---|
vpc |
./modules/vpc |
VPC with public/private subnets, IGW, NAT, security groups |
rds_aurora_postgresql |
./modules/rds-aurora-postgresql |
Aurora PostgreSQL with encryption, backups, monitoring |
eks_karpenter |
./modules/eks-karpenter |
EKS cluster + Karpenter for node autoscaling |
eks_addons |
./modules/eks-addons |
Karpenter Helm chart, AWS LB Controller |
cognito |
./modules/cognito |
Cognito User Pool & App Client (callback URLs auto-derived from frontend_domain) |
waf |
./modules/waf |
Web Application Firewall (auto-enabled after ALBs are ready) |
global_accelerator |
./modules/global-accelerator |
Optional GA for global edge routing |
| Setting | Production | Non-Production |
|---|---|---|
deletion_protection |
true |
false |
backup_retention_period |
7 days | 1 day |
performance_insights |
Enabled | Disabled |
monitoring_interval |
60s | 0 (disabled) |
skip_final_snapshot |
false |
true |
apply_immediately |
false |
true |
flow_logs (GA) |
Enabled | Disabled |
The system supports two OAuth authentication methods: AWS Cognito and Microsoft Entra ID (selectable during deploy-all.sh --step 0). There is no local username/password authentication. The admin dashboard uses JWT (access + refresh tokens), while the gateway APIs use API keys (kbr_ prefix) -- via Authorization: Bearer for OpenAI-compatible endpoints or x-api-key header for Anthropic endpoints. Both auth methods validate the same kbr_ tokens. Cognito callback URLs are automatically derived from frontend_domain in terraform.tfvars.
sequenceDiagram
participant User as User (Browser)
participant FE as Frontend
participant BE as Backend
participant IdP as Identity Provider<br/>(Microsoft / Cognito)
participant DB as PostgreSQL
User->>FE: Click "Sign in with Microsoft"
FE->>BE: GET /admin/auth/microsoft/login?redirect_uri=...
BE->>BE: Generate PKCE code_verifier + code_challenge (S256)
BE->>DB: Store OAuthState (state, code_verifier, 10 min TTL)
BE-->>FE: {authorization_url (includes code_challenge), state}
FE->>IdP: Redirect to authorization_url
User->>IdP: Authenticate + consent
IdP-->>FE: Redirect to /auth/microsoft/callback?code=...&state=...
FE->>BE: POST /admin/auth/microsoft/callback {code, state, redirect_uri}
BE->>DB: Verify OAuthState (CSRF check), retrieve code_verifier
BE->>IdP: Exchange code for tokens (with code_verifier)
IdP-->>BE: {access_token, id_token}
BE->>IdP: Fetch user profile (MS Graph / Cognito userInfo)
IdP-->>BE: {email, name, sub}
BE->>DB: Find or create User (auth_method=MICROSOFT)
BE->>DB: Create RefreshToken + AuditLog
BE-->>FE: {access_token, user} + Set-Cookie: kbr_refresh_token (HttpOnly)
FE->>FE: Store access_token in localStorage, redirect to dashboard
The same kbr_ API keys work for both OpenAI-compatible and Anthropic endpoints. The only difference is the header format:
- OpenAI path:
Authorization: Bearer kbr_xxx(extracted from Bearer token) - Anthropic path:
x-api-key: kbr_xxx(extracted fromx-api-keyheader)
Both paths use the same validation logic (Redis cache → DB fallback).
sequenceDiagram
participant Client as API Client<br/>(OpenAI or Anthropic SDK)
participant BE as Backend (FastAPI)
participant Redis as Redis (optional)
participant DB as PostgreSQL
alt OpenAI-compatible endpoint
Client->>BE: POST /v1/chat/completions<br/>Authorization: Bearer kbr_xxx...
BE->>BE: Extract token from Bearer header
else Anthropic endpoint
Client->>BE: POST /v1/messages<br/>x-api-key: kbr_xxx...
BE->>BE: Extract token from x-api-key header
end
alt Redis available
BE->>Redis: Check token cache (SHA256 hash)
Redis-->>BE: Cache hit / miss
end
alt Cache miss or Redis unavailable
BE->>DB: SELECT * FROM api_tokens<br/>WHERE token_hash = SHA256(kbr_xxx)
DB-->>BE: APIToken record
end
BE->>BE: Check is_active, is_expired
BE->>DB: SUM(cost_usd) for token
BE->>BE: Check quota_usd >= total_used
BE-->>Client: Token validated (or 401/429 error)
| Property | JWT Access Token | JWT Refresh Token | API Key (kbr_) |
|---|---|---|---|
| Lifetime | 30 minutes | 7 days | Until expiry or revocation |
| Used by | Admin dashboard | Admin dashboard (refresh) | OpenAI / Anthropic clients |
| Storage | localStorage | HttpOnly cookie (kbr_refresh_token, Path=/admin/auth) |
Client configuration |
| Validation | JWT decode + signature | DB lookup (hash + family) | DB lookup (SHA256 hash) |
| Rotation | On refresh | On each use (new token issued) | Manual |
This section details the full lifecycle of gateway requests. The proxy supports these API paths:
- OpenAI path → AWS Bedrock (
/v1/chat/completions, non-Gemini / non-mantle models): Full translation between OpenAI and Bedrock formats - OpenAI path → Google Gemini (
/v1/chat/completions,gemini-*models):GeminiClientconverts OpenAI ↔ Gemini native format;chat.pyis format-agnostic - OpenAI path → AWS mantle (
/v1/chat/completionsand/v1/messages,openai.gpt-5.5/openai.gpt-5.4):is_openai_mantle_model()routes the request toMantleClient, which converts OpenAI ChatCompletions ↔ OpenAI Responses (lossy) and calls mantle'sPOST /responsesover SigV4 - mantle native passthrough (
/v1/responses): zero-conversion OpenAI Responses API forwarded directly to mantle - Anthropic path (
/v1/messages): Translated (not a true passthrough). The format is close — Bedrock InvokeModel natively uses Anthropic Messages API format — but the request is still mapped throughto_bedrock→BedrockRequest, which has nocache_controlfield, so inline cache breakpoints are dropped (the proxy injects its own).to_bedrock_with_passthroughexists but is not currently wired in.
OpenAI GPT-5.5/5.4 can therefore be reached via three access paths: /v1/chat/completions (lossy conversion), /v1/messages (lossy conversion), and /v1/responses (native, zero conversion).
sequenceDiagram
participant Client as OpenAI Client
participant MW as Middleware Stack
participant Chat as chat.py endpoint
participant Deps as deps.py (DI)
participant TokenSvc as TokenService
participant DB as PostgreSQL
participant Translator as RequestTranslator /<br/>ResponseTranslator
participant Bedrock as BedrockClient<br/>(semaphore=50)
participant AWS as AWS Bedrock API
participant Pricing as ModelPricing
participant BG as BackgroundTaskManager
Client->>MW: POST /v1/chat/completions<br/>Authorization: Bearer kbr_xxx
MW->>MW: SecurityMiddleware: origin + CSRF check
MW->>MW: CORS headers
MW->>Chat: Route to handler
Chat->>Deps: Depends(get_current_token)
Deps->>TokenSvc: validate_token(kbr_xxx)
TokenSvc->>DB: SELECT by token_hash
DB-->>TokenSvc: APIToken
TokenSvc-->>Chat: Validated APIToken
Chat->>DB: SUM(cost_usd) WHERE token_id = ?
DB-->>Chat: total_used
Chat->>Chat: Check quota: total_used < quota_usd
Note over Chat: 429 if quota exceeded
Chat->>DB: SELECT models WHERE token_id = ?<br/>AND is_active AND NOT is_deleted
DB-->>Chat: Allowed model names
Chat->>Chat: Normalize request.model + allowed_models<br/>(strip geo prefix + version suffix)
Chat->>Chat: Check normalized model in normalized allowed set
Note over Chat: 403 if model not allowed<br/>On match: replace request.model with Bedrock ID from DB
Chat->>Translator: RequestTranslator.openai_to_bedrock(request)
Translator-->>Chat: BedrockRequest
alt Streaming (stream=true)
Chat->>Bedrock: invoke_stream(model, request)
Note over Bedrock: Acquire semaphore (50 max)
alt Anthropic model
Bedrock->>AWS: invoke_model_with_response_stream(body)
loop Stream events
AWS-->>Bedrock: Anthropic SSE event (JSON)
Bedrock->>Bedrock: _anthropic_event_to_bedrock()
end
else Non-Anthropic model (Nova, DeepSeek, etc.)
Bedrock->>AWS: converse_stream(params)
loop Stream events
AWS-->>Bedrock: Converse stream event
Bedrock->>Bedrock: _converse_stream_event_to_bedrock()
end
end
Bedrock-->>Chat: BedrockStreamEvent
Chat->>Translator: create_stream_chunk(...)
Translator-->>Chat: SSE formatted chunk
Chat-->>Client: data: {...}\n\n
Chat-->>Client: data: [DONE]\n\n
Note over Bedrock: Release semaphore
else Non-streaming
Chat->>Bedrock: invoke(model, request)
Note over Bedrock: Acquire semaphore
alt Anthropic model
Bedrock->>AWS: invoke_model(body)
AWS-->>Bedrock: Anthropic Messages API JSON
else Non-Anthropic model (Nova, DeepSeek, etc.)
Bedrock->>AWS: converse(params)
AWS-->>Bedrock: Converse API JSON
end
Note over Bedrock: Release semaphore
Bedrock-->>Chat: BedrockResponse
Chat->>Translator: bedrock_to_openai(response)
Translator-->>Chat: ChatCompletionResponse
Chat-->>Client: JSON response
end
Chat->>BG: background: record_usage(...)
BG->>Pricing: calculate_cost(model, tokens)
Pricing-->>BG: cost_usd
BG->>DB: INSERT INTO usage_records
When the requested model starts with gemini-, chat.py routes to _handle_gemini_request(). The GeminiClient handles all format conversion internally; chat.py always receives an OpenAI-format response dict.
sequenceDiagram
participant Client as OpenAI Client
participant Chat as chat.py endpoint
participant GC as GeminiClient
participant Gemini as Google Gemini API<br/>(generativelanguage.googleapis.com)
Client->>Chat: POST /v1/chat/completions<br/>model: "gemini-2.5-flash"
Chat->>Chat: is_gemini_model() → true
Chat->>Chat: Quota + model access check
alt Non-streaming (stream=false, or image model)
Chat->>GC: invoke(payload, api_key)
GC->>GC: _openai_to_gemini_payload()<br/>messages→contents+systemInstruction<br/>max_tokens→maxOutputTokens, tools→functionDeclarations
GC->>Gemini: POST /v1beta/models/{model}:generateContent
Gemini-->>GC: GenerateContentResponse (JSON)
GC->>GC: _gemini_response_to_openai()<br/>candidates→choices, usageMetadata→usage<br/>inlineData→image_url (base64)
GC-->>Chat: OpenAI-format dict
Chat-->>Client: JSON response
else Streaming (stream=true)
Chat->>GC: invoke_stream(payload, api_key)
GC->>GC: _openai_to_gemini_payload()
GC->>Gemini: POST /v1beta/models/{model}:streamGenerateContent?alt=sse
loop SSE chunks
Gemini-->>GC: data: {GenerateContentResponse chunk}
GC->>GC: _gemini_chunk_to_sse()<br/>→ OpenAI chat.completion.chunk format
GC-->>Chat: OpenAI SSE string
end
GC-->>Chat: data: [DONE]
Chat-->>Client: SSE stream
end
Chat->>Chat: background: record_usage(cached_tokens extracted from usage)
Key conversion mappings (OpenAI → Gemini):
| OpenAI field | Gemini field |
|---|---|
messages[role=system] |
systemInstruction.parts |
messages[role=user/assistant] |
contents[role=user/model] |
messages[role=tool] |
contents[role=user].parts[functionResponse] |
tool_calls in assistant message |
parts[functionCall] |
max_tokens |
generationConfig.maxOutputTokens |
temperature |
generationConfig.temperature |
top_p |
generationConfig.topP |
stop |
generationConfig.stopSequences |
tools[].function |
tools[].functionDeclarations[] |
tool_choice: "none/auto/required" |
toolConfig.functionCallingConfig.mode: NONE/AUTO/ANY |
image_url (base64 data URI) |
inlineData.mimeType + inlineData.data |
Key conversion mappings (Gemini → OpenAI):
| Gemini field | OpenAI field |
|---|---|
candidates[0].content.parts[text] |
choices[0].message.content (string) |
candidates[0].content.parts[functionCall] |
choices[0].message.tool_calls[] |
candidates[0].content.parts[inlineData] |
choices[0].message.content (array with image_url) |
candidates[0].finishReason: STOP |
choices[0].finish_reason: stop |
candidates[0].finishReason: MAX_TOKENS |
choices[0].finish_reason: length |
candidates[0].finishReason: SAFETY |
choices[0].finish_reason: content_filter |
usageMetadata.promptTokenCount |
usage.prompt_tokens |
usageMetadata.candidatesTokenCount |
usage.completion_tokens |
usageMetadata.cachedContentTokenCount |
usage.prompt_tokens_details.cached_tokens |
When the requested model matches openai.gpt-5.5 / openai.gpt-5.4 (exact match via is_openai_mantle_model()), chat.py (and messages.py) routes to MantleClient before BedrockClient is touched. MantleClient SigV4-signs the request (service bedrock, reusing the EKS Pod IRSA credential chain incl. SessionToken), resolves the region via resolve_mantle_region() (GPT-5.5 → us-east-2 only; GPT-5.4 → us-east-2 + us-west-2), and calls mantle's OpenAI Responses API. The /v1/responses endpoint forwards natively with zero conversion.
sequenceDiagram
participant Client as OpenAI Client
participant Chat as chat.py / responses.py
participant MC as MantleClient
participant Mantle as AWS mantle<br/>(bedrock-mantle.{region}.api.aws)
Client->>Chat: POST /v1/chat/completions<br/>model: "openai.gpt-5.5"
Chat->>Chat: is_openai_mantle_model() → true
Chat->>Chat: Quota + model access check
Chat->>MC: resolve_mantle_region(model)
alt ChatCompletions conversion (/v1/chat/completions, /v1/messages)
Chat->>MC: invoke / invoke_stream(payload)
MC->>MC: _openai_to_responses()<br/>system→developer role, max_tokens→max_output_tokens<br/>image→input_image, response_format→text.format<br/>reasoning.effort
MC->>MC: SigV4 sign (SigV4Auth/AWSRequest, service=bedrock)
MC->>Mantle: POST /openai/v1/responses (stream:true)
loop Named SSE events
Mantle-->>MC: response.output_text.delta / response.completed
MC->>MC: _responses_to_openai()<br/>→ chat.completion(.chunk) format
MC-->>Chat: OpenAI dict / SSE string
end
Chat-->>Client: JSON / SSE stream
else Native passthrough (/v1/responses)
Chat->>MC: responses_passthrough / responses_passthrough_stream(body)
MC->>MC: SigV4 sign (no conversion)
MC->>Mantle: POST /openai/v1/responses
Mantle-->>MC: Responses payload / named SSE events
MC-->>Chat: raw Responses payload / SSE
Chat-->>Client: native Responses response
end
Chat->>Chat: background: record_usage(...)
The Anthropic path is close to the native format — Bedrock's InvokeModel API accepts Anthropic Messages API format — but it is still a translation (to_bedrock → BedrockRequest), not a byte passthrough: inline cache_control markers are dropped (see Prompt Caching). Key differences from the OpenAI path:
- Auth via
x-api-keyheader instead ofAuthorization: Bearer - Thinking blocks are preserved in responses (OpenAI path skips them)
- Streaming uses Anthropic SSE format (
event: type\ndata: {json}\n\n) instead of OpenAI format (data: {json}\n\n)
sequenceDiagram
participant Client as Anthropic Client
participant MW as Middleware Stack
participant Msg as messages.py endpoint
participant Deps as deps.py (DI)
participant TokenSvc as TokenService
participant DB as PostgreSQL
participant Translator as AnthropicRequestTranslator /<br/>AnthropicResponseTranslator
participant Bedrock as BedrockClient<br/>(semaphore=50)
participant AWS as AWS Bedrock API
participant BG as BackgroundTaskManager
Client->>MW: POST /v1/messages<br/>x-api-key: kbr_xxx
MW->>MW: SecurityMiddleware + CORS
MW->>Msg: Route to handler
Msg->>Deps: Depends(get_current_token_from_api_key)
Deps->>TokenSvc: validate_token(kbr_xxx)
TokenSvc->>DB: SELECT by token_hash
DB-->>TokenSvc: APIToken
TokenSvc-->>Msg: Validated APIToken
Msg->>Msg: Quota check + model access check<br/>(normalizes Anthropic short names to Bedrock IDs)
Msg->>Translator: to_bedrock(request)
Note over Translator: Maps to BedrockRequest;<br/>inline cache_control is DROPPED<br/>(schema has no such field)
Translator-->>Msg: BedrockRequest
alt Streaming (stream=true)
Msg->>Bedrock: invoke_stream(model, request)
Bedrock->>AWS: invoke_model_with_response_stream(body)
loop Stream events
AWS-->>Bedrock: Anthropic SSE event
Bedrock-->>Msg: BedrockStreamEvent
end
Msg->>Translator: bedrock_stream_to_anthropic_events(event)
Translator-->>Msg: Anthropic SSE formatted string
Msg-->>Client: event: content_block_delta\ndata: {...}\n\n
Msg-->>Client: event: message_stop\ndata: {...}\n\n
else Non-streaming
Msg->>Bedrock: invoke(model, request)
Bedrock->>AWS: invoke_model(body)
AWS-->>Bedrock: Anthropic Messages API JSON
Bedrock-->>Msg: BedrockResponse
Msg->>Translator: bedrock_to_anthropic(response)
Translator-->>Msg: AnthropicMessagesResponse
Msg-->>Client: JSON response
end
Msg->>BG: background: record_usage(...)
| Path | Translation | Notes |
|---|---|---|
| OpenAI → Bedrock | Three-phase (OpenAI → BedrockRequest → Bedrock API → BedrockResponse → OpenAI) |
Full translation; tool calls, images, Bedrock extensions |
| OpenAI → Gemini | GeminiClient internal conversion (OpenAI → Gemini GenerateContentRequest / back) |
Native generateContent API; no OpenAI-compat layer |
| OpenAI → mantle | MantleClient internal conversion (OpenAI ChatCompletions ↔ OpenAI Responses) |
SigV4 to POST /responses; lossy. Native /v1/responses is zero-conversion passthrough |
| Anthropic → Bedrock | Translated via to_bedrock → BedrockRequest (invoke_model) |
Format is close to native, but it is a translation, not a passthrough: thinking blocks are preserved; inline cache_control markers are dropped (BedrockRequest has no such field). The proxy injects its own breakpoints. to_bedrock_with_passthrough exists but is not wired in. |
For non-Anthropic Bedrock models (Nova, DeepSeek, Mistral, Llama, etc.), the Bedrock path uses the Converse API via converse/converse_stream.
For the complete translation pipeline documentation, see Request Translation.
Cost is calculated per request based on actual AWS Bedrock pricing for each model. The ModelPricing class in backend/app/services/pricing.py fetches per-token rates from the database. Pricing region is determined dynamically via BedrockClient.resolve_model(), which uses the inference profile cache to identify the actual region where the model runs.
graph TD
A["API Response received<br/>with token counts"] --> B["Extract model, prompt_tokens,<br/>completion_tokens,<br/>cache_creation_input_tokens,<br/>cache_read_input_tokens"]
B --> C{"Model in<br/>PRICING table?"}
C -->|"Yes"| D["Get model-specific<br/>(input_rate, output_rate)"]
C -->|"No"| E["Fallback: Claude 3.5<br/>Sonnet pricing"]
D --> F["cost = prompt_tokens × input_rate<br/>+ completion_tokens × output_rate<br/>+ cache_write_tokens × input_rate × 1.25<br/>+ cache_read_tokens × input_rate × 0.1"]
E --> F
F --> G["INSERT UsageRecord<br/>(cost_usd, tokens, cache tokens, model)"]
G --> H["Token quota check on<br/>next request:<br/>SUM(cost_usd) vs quota_usd"]
Prompt Cache Pricing: When prompt caching is enabled, cache write tokens are charged at 1.25x and cache read tokens at 0.1x the base input price. See Dynamic Pricing System for details.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Typical Use Case |
|---|---|---|---|
| Claude 3.5 Sonnet v2 | $3.00 | $15.00 | Balanced performance |
| Claude 3.5 Sonnet | $3.00 | $15.00 | Balanced performance |
| Claude 3 Sonnet | $3.00 | $15.00 | Standard tasks |
| Claude 3 Haiku | $0.25 | $1.25 | Fast, cost-effective |
| Claude 3 Opus | $15.00 | $75.00 | Highest intelligence |
| Mistral Large | $0.50 | $1.50 | European alternative |
| Mistral Small | $1.00 | $3.00 | Lightweight tasks |
| Llama 3 70B | $2.65 | $3.50 | Open source, large |
| Llama 3 8B | $0.30 | $0.60 | Open source, small |
Request: Claude 3 Haiku, 10,000 input tokens, 5,000 output tokens
input_cost = 10,000 * ($0.25 / 1,000,000) = $0.0025
output_cost = 5,000 * ($1.25 / 1,000,000) = $0.00625
total_cost = $0.0025 + $0.00625 = $0.00875
Each API token (APIToken.quota_usd) can have an optional spending limit. The quota check happens at the beginning of each request:
- Query
SUM(cost_usd)fromusage_recordsfor the token - Compare against
quota_usd - If
total_used >= quota_usd, return HTTP 429 with message:Token quota exceeded. Used: $X.XX, Quota: $Y.YY
Usage recording is performed asynchronously via BackgroundTaskManager to avoid blocking the response to the client. The cost is calculated using ModelPricing.calculate_cost() with a fallback to Claude 3.5 Sonnet pricing if the model is not found in the pricing table.
Displayed costs are estimates based on token usage. Actual AWS billing may differ due to pricing updates, regional variations, additional AWS fees, and rounding differences in token counting.
| Document | Description |
|---|---|
| Request Translation | Full request/response translation pipeline (OpenAI → Bedrock → Anthropic) |
| Dynamic Pricing System | Price fetching, cache-aware cost calculation, and pricing table display |
| API Reference | Complete endpoint documentation with request/response examples |
| OAuth Setup | Microsoft and Cognito OAuth configuration |
| Deployment | Production and non-production deployment guide |