Skip to content

Commit 5c0c9e4

Browse files
authored
feat(providers): add OCI Generative AI example provider profile (#3904)
Add an example oci-genai inference profile for Oracle Cloud Infrastructure Generative AI through its OpenAI-compatible endpoint. The profile injects a compartment-scoped Generative AI API key as a bearer token only at the regional OCI inference hosts and only under /openai/v1 with GET, POST, and DELETE, so the sandbox never holds the key and the key cannot reach any other OCI surface. The header comments carry the OCI-side setup (create the IAM policy before the key, least-privilege statement, key creation and rotation), the operations verified through the sandbox proxy with a real key (chat completions with streaming, tool calling, vision input, embeddings, and the Responses API), the OCI error messages operators will meet, and the realm and signed-transport caveats. Scoped to the profile YAML per #3906; the only code change is the entry in the profile listing test, which enumerates providers/*.yaml. Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com>
1 parent ba16b9f commit 5c0c9e4

2 files changed

Lines changed: 94 additions & 0 deletions

File tree

‎crates/openshell-server/src/grpc/provider.rs‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6301,6 +6301,7 @@ mod tests {
63016301
"google-cloud",
63026302
"google-vertex-ai",
63036303
"nvidia",
6304+
"oci-genai",
63046305
"openai",
63056306
"openrouter",
63066307
"pypi"

‎providers/oci-genai.yaml‎

Lines changed: 93 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,93 @@
1+
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
2+
# SPDX-License-Identifier: Apache-2.0
3+
4+
# Example provider profile. OpenShell does not load it; import it explicitly:
5+
# openshell provider profile lint -f providers/oci-genai.yaml
6+
# openshell provider profile import -f providers/oci-genai.yaml --global
7+
#
8+
# Copy and edit this file rather than importing it unchanged. `binaries` is the
9+
# least-privilege control that decides which processes may reach the endpoints
10+
# below, so it has to name the paths in *your* image.
11+
#
12+
# Client binaries: curl. Add the interpreter or agent CLI that calls the
13+
# OpenAI-compatible endpoint (python, node, ...) before
14+
# relying on binary-scoped attribution for this profile.
15+
# Reference layout: any image with curl or an OpenAI SDK. Point the SDK at
16+
# https://inference.generativeai.<region>.oci.oraclecloud.com/openai/v1
17+
# and pass OCI_GENAI_API_KEY as the API key; the sandbox
18+
# proxy resolves the placeholder only at the hosts below.
19+
# Credential scope: one OCI Generative AI API key (sk-...). The key is scoped
20+
# to the compartment it was created in, so requests need no
21+
# opc-compartment-id header and the sandbox never learns a
22+
# compartment OCID.
23+
# Endpoint access: the regional OCI Generative AI inference hosts in the
24+
# commercial realm (oraclecloud.com), TLS terminated, L7
25+
# enforced, limited to the OpenAI-compatible /openai/v1
26+
# surface with GET, POST, and DELETE. Generative AI API keys
27+
# are only valid there; the native /20231130 API needs OCI
28+
# request signing, which this profile does not grant.
29+
# Smoke test: openshell sandbox create --provider <name> -- \
30+
# curl -sS --compressed \
31+
# https://inference.generativeai.us-chicago-1.oci.oraclecloud.com/openai/v1/chat/completions \
32+
# -H "Authorization: Bearer $OCI_GENAI_API_KEY" \
33+
# -H 'Content-Type: application/json' \
34+
# -d '{"model":"meta.llama-3.3-70b-instruct","messages":[{"role":"user","content":"Reply with OK"}]}'
35+
#
36+
# OCI-side setup, in this order:
37+
# 1. Create the IAM policy BEFORE the key. A key minted before its policy can
38+
# stay unauthorized long after the policy lands, and OCI returns the same
39+
# 401 for an unknown key and an unauthorized one. Least privilege:
40+
# allow any-user to use generative-ai-family in compartment id <ocid>
41+
# where ALL {request.principal.type='generativeaiapikey'}
42+
# Pin it to one key later by adding request.principal.id='<api-key-ocid>'.
43+
# 2. Create the key in that compartment and region, with an expiry:
44+
# oci generative-ai api-key create --compartment-id <ocid> \
45+
# --region us-chicago-1 --display-name openshell-agents \
46+
# --key-details '[{"keyName":"primary","timeExpiry":"2027-01-01T00:00:00Z"}]'
47+
# The sk-... secret is returned once. Rotate with `api-key renew`,
48+
# disable with `api-key set-api-key-state`; sandboxes are unaffected
49+
# because the gateway, not the sandbox, holds the value.
50+
# 3. openshell provider create --name my-oci-genai --type oci-genai \
51+
# --credential OCI_GENAI_API_KEY="$OCI_GENAI_API_KEY"
52+
#
53+
# Verified through the sandbox proxy with this profile: chat completions with
54+
# stream=true, tool calling, vision input via data: URLs (for example
55+
# meta.llama-4-maverick-17b-128e-instruct-fp8 and google.gemini-2.5-flash,
56+
# including a 263 KB request body), embeddings with openai.text-embedding-3-small,
57+
# and the Responses API with stream=true. Model ids use the OCI form, such as
58+
# meta.llama-3.3-70b-instruct or openai.gpt-oss-120b. Expect from OCI, not from
59+
# the proxy: 404 "Entity with key <model> not found" for models not served on
60+
# this endpoint, 404 on GET /openai/v1/models (not a health check), and
61+
# 400 "Unsupported OpenAI operation" for Cohere embeddings.
62+
#
63+
# Other realms use their realm domain instead of oraclecloud.com: copy the
64+
# profile and add a matching endpoint entry. Signed OCI transports (API key
65+
# signing, security token, instance/resource principal, OKE workload identity)
66+
# need proxy-side request signing; see https://github.com/NVIDIA/OpenShell/issues/3879.
67+
68+
id: oci-genai
69+
display_name: OCI Generative AI
70+
description: OCI Generative AI inference through the OpenAI-compatible endpoint with a compartment-scoped Generative AI API key
71+
category: inference
72+
inference_capable: true
73+
credentials:
74+
- name: api_key
75+
description: OCI Generative AI API key (compartment-scoped, sent as a bearer token)
76+
env_vars: [OCI_GENAI_API_KEY]
77+
required: true
78+
auth_style: bearer
79+
header_name: authorization
80+
discovery:
81+
credentials: [api_key]
82+
endpoints:
83+
# `*` matches exactly one DNS label, which is the region identifier, for
84+
# example us-chicago-1, eu-frankfurt-1, or ap-osaka-1.
85+
- host: "inference.generativeai.*.oci.oraclecloud.com"
86+
port: 443
87+
protocol: rest
88+
enforcement: enforce
89+
rules:
90+
- allow: { method: POST, path: "/openai/v1/**" }
91+
- allow: { method: GET, path: "/openai/v1/**" }
92+
- allow: { method: DELETE, path: "/openai/v1/**" }
93+
binaries: [/usr/bin/curl, /usr/local/bin/curl]

0 commit comments

Comments
 (0)