Enable Prompt Caching on amazon bedrock Claude models using Autogen 0.4.7 #6504
Replies: 2 comments
|
Hello Eric Zhu (@ekzhu) , could you please help me on that? Many thanks in advance |
|
Thanks for pasting the full client setup and the blog link, that made this easy to reproduce end to end. Short answer: nothing in that stack exposes prompt caching today, so there is no setting to flip. SK's Bedrock connector serialises the system prompt as Because your setup already owns a boto3 client, a botocore event hook can append the sk_client = BedrockChatCompletion(model_id=MODEL)
def add_cache_point(params, **kwargs):
if params.get("system"): # Converse wants a content block before the cache point
params["system"].append({"cachePoint": {"type": "default"}})
for op in ("Converse", "ConverseStream"):
sk_client.bedrock_runtime_client.meta.events.register(f"before-parameter-build.bedrock-runtime.{op}", add_cache_point)
# optional: the cache counters only exist in the raw Converse response
sk_client.bedrock_runtime_client.meta.events.register(
"after-call.bedrock-runtime.Converse", lambda parsed, **kw: print("converse usage:", parsed["usage"]))
client = SKChatCompletionAdapter(sk_client, kernel=Kernel(), prompt_settings=..., model_info=...)Measured on autogen
Two things to watch: Where this lives, your model IDs, and the native clientWhy there is no setting. Your three model IDs. On my account the 3.5 Sonnet and 3.7 Sonnet IDs now return
client = AnthropicBedrockChatCompletionClient(model=MODEL, bedrock_info=BedrockInfo(aws_region="us-east-1"), model_info=...)
_create = client._client.messages.create
async def create_with_cache(**kw):
if isinstance(kw.get("system"), str):
kw["system"] = [{"type": "text", "text": kw["system"], "cache_control": {"type": "ephemeral"}}]
return await _create(**kw)
client._client.messages.create = create_with_cacheThat first call already read the entry written by the SK run above: same model, same prompt, within the TTL. |
Uh oh!
There was an error while loading. Please reload this page.
Hello experts,
I realised that Amazon implement Prompt caching on Claude models as they showed on the next blog: https://aws.amazon.com/blogs/machine-learning/effectively-use-prompt-caching-on-amazon-bedrock/
I would like to leverage that feature using Autogen 0.4.7.
Mi model client is configured using the follow code:
`
from autogen_ext.models.semantic_kernel import SKChatCompletionAdapter
from semantic_kernel import Kernel
from semantic_kernel.connectors.ai.bedrock import BedrockChatCompletion, BedrockPromptExecutionSettings
from semantic_kernel.memory.null_memory import NullMemory
import os
from enum import Enum
class LLMModel(str, Enum):
CLAUDE_3_5_SONNET = "eu.anthropic.claude-3-5-sonnet-20240620-v1:0"
CLAUDE_3_HAIKU = "anthropic.claude-3-haiku-20240307-v1:0"
CLAUDE_3_7_SONNET = "us.anthropic.claude-3-7-sonnet-20250219-v1:0"
def get_bedrock_model_client(
model_id=LLMModel.CLAUDE_3_5_SONNET,
region="us-east-1",
temperature=0.0,
price=[0.003, 0.015]
):
"""
model_id: eu.anthropic.claude-3-5-sonnet-20240620-v1:0 / anthropic.claude-3-haiku-20240307-v1:0
"""
os.environ["AWS_DEFAULT_REGION"] = region
`
I tried to use "cache_seed" but it doesn't work.
Could you please tell me how can I activate Prompt Caching on AWS bedrock Claude models, using Autogen 0.4.7?
Many thanks in advance.
Regards,,
Alan
All reactions