|
| 1 | +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. |
| 2 | +# SPDX-License-Identifier: Apache-2.0 |
| 3 | + |
| 4 | +# Example provider profile. OpenShell does not load it; import it explicitly: |
| 5 | +# openshell provider profile lint -f providers/oci-genai.yaml |
| 6 | +# openshell provider profile import -f providers/oci-genai.yaml --global |
| 7 | +# |
| 8 | +# Copy and edit this file rather than importing it unchanged. `binaries` is the |
| 9 | +# least-privilege control that decides which processes may reach the endpoints |
| 10 | +# below, so it has to name the paths in *your* image. |
| 11 | +# |
| 12 | +# Client binaries: curl. Add the interpreter or agent CLI that calls the |
| 13 | +# OpenAI-compatible endpoint (python, node, ...) before |
| 14 | +# relying on binary-scoped attribution for this profile. |
| 15 | +# Reference layout: any image with curl or an OpenAI SDK. Point the SDK at |
| 16 | +# https://inference.generativeai.<region>.oci.oraclecloud.com/openai/v1 |
| 17 | +# and pass OCI_GENAI_API_KEY as the API key; the sandbox |
| 18 | +# proxy resolves the placeholder only at the hosts below. |
| 19 | +# Credential scope: one OCI Generative AI API key (sk-...). The key is scoped |
| 20 | +# to the compartment it was created in, so requests need no |
| 21 | +# opc-compartment-id header and the sandbox never learns a |
| 22 | +# compartment OCID. |
| 23 | +# Endpoint access: the regional OCI Generative AI inference hosts in the |
| 24 | +# commercial realm (oraclecloud.com), TLS terminated, L7 |
| 25 | +# enforced, limited to the OpenAI-compatible /openai/v1 |
| 26 | +# surface with GET, POST, and DELETE. Generative AI API keys |
| 27 | +# are only valid there; the native /20231130 API needs OCI |
| 28 | +# request signing, which this profile does not grant. |
| 29 | +# Smoke test: openshell sandbox create --provider <name> -- \ |
| 30 | +# curl -sS --compressed \ |
| 31 | +# https://inference.generativeai.us-chicago-1.oci.oraclecloud.com/openai/v1/chat/completions \ |
| 32 | +# -H "Authorization: Bearer $OCI_GENAI_API_KEY" \ |
| 33 | +# -H 'Content-Type: application/json' \ |
| 34 | +# -d '{"model":"meta.llama-3.3-70b-instruct","messages":[{"role":"user","content":"Reply with OK"}]}' |
| 35 | +# |
| 36 | +# OCI-side setup, in this order: |
| 37 | +# 1. Create the IAM policy BEFORE the key. A key minted before its policy can |
| 38 | +# stay unauthorized long after the policy lands, and OCI returns the same |
| 39 | +# 401 for an unknown key and an unauthorized one. Least privilege: |
| 40 | +# allow any-user to use generative-ai-family in compartment id <ocid> |
| 41 | +# where ALL {request.principal.type='generativeaiapikey'} |
| 42 | +# Pin it to one key later by adding request.principal.id='<api-key-ocid>'. |
| 43 | +# 2. Create the key in that compartment and region, with an expiry: |
| 44 | +# oci generative-ai api-key create --compartment-id <ocid> \ |
| 45 | +# --region us-chicago-1 --display-name openshell-agents \ |
| 46 | +# --key-details '[{"keyName":"primary","timeExpiry":"2027-01-01T00:00:00Z"}]' |
| 47 | +# The sk-... secret is returned once. Rotate with `api-key renew`, |
| 48 | +# disable with `api-key set-api-key-state`; sandboxes are unaffected |
| 49 | +# because the gateway, not the sandbox, holds the value. |
| 50 | +# 3. openshell provider create --name my-oci-genai --type oci-genai \ |
| 51 | +# --credential OCI_GENAI_API_KEY="$OCI_GENAI_API_KEY" |
| 52 | +# |
| 53 | +# Verified through the sandbox proxy with this profile: chat completions with |
| 54 | +# stream=true, tool calling, vision input via data: URLs (for example |
| 55 | +# meta.llama-4-maverick-17b-128e-instruct-fp8 and google.gemini-2.5-flash, |
| 56 | +# including a 263 KB request body), embeddings with openai.text-embedding-3-small, |
| 57 | +# and the Responses API with stream=true. Model ids use the OCI form, such as |
| 58 | +# meta.llama-3.3-70b-instruct or openai.gpt-oss-120b. Expect from OCI, not from |
| 59 | +# the proxy: 404 "Entity with key <model> not found" for models not served on |
| 60 | +# this endpoint, 404 on GET /openai/v1/models (not a health check), and |
| 61 | +# 400 "Unsupported OpenAI operation" for Cohere embeddings. |
| 62 | +# |
| 63 | +# Other realms use their realm domain instead of oraclecloud.com: copy the |
| 64 | +# profile and add a matching endpoint entry. Signed OCI transports (API key |
| 65 | +# signing, security token, instance/resource principal, OKE workload identity) |
| 66 | +# need proxy-side request signing; see https://github.com/NVIDIA/OpenShell/issues/3879. |
| 67 | + |
| 68 | +id: oci-genai |
| 69 | +display_name: OCI Generative AI |
| 70 | +description: OCI Generative AI inference through the OpenAI-compatible endpoint with a compartment-scoped Generative AI API key |
| 71 | +category: inference |
| 72 | +inference_capable: true |
| 73 | +credentials: |
| 74 | + - name: api_key |
| 75 | + description: OCI Generative AI API key (compartment-scoped, sent as a bearer token) |
| 76 | + env_vars: [OCI_GENAI_API_KEY] |
| 77 | + required: true |
| 78 | + auth_style: bearer |
| 79 | + header_name: authorization |
| 80 | +discovery: |
| 81 | + credentials: [api_key] |
| 82 | +endpoints: |
| 83 | + # `*` matches exactly one DNS label, which is the region identifier, for |
| 84 | + # example us-chicago-1, eu-frankfurt-1, or ap-osaka-1. |
| 85 | + - host: "inference.generativeai.*.oci.oraclecloud.com" |
| 86 | + port: 443 |
| 87 | + protocol: rest |
| 88 | + enforcement: enforce |
| 89 | + rules: |
| 90 | + - allow: { method: POST, path: "/openai/v1/**" } |
| 91 | + - allow: { method: GET, path: "/openai/v1/**" } |
| 92 | + - allow: { method: DELETE, path: "/openai/v1/**" } |
| 93 | +binaries: [/usr/bin/curl, /usr/local/bin/curl] |
0 commit comments