[{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/tags/agents/","section":"Tags","summary":"","title":"Agents","type":"tags"},{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/categories/ai/","section":"Categories","summary":"","title":"AI","type":"categories"},{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/tags/azure/","section":"Tags","summary":"","title":"Azure","type":"tags"},{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/","section":"BeyondElastic","summary":"","title":"BeyondElastic","type":"page"},{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/categories/","section":"Categories","summary":"","title":"Categories","type":"categories"},{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/tags/foundry/","section":"Tags","summary":"","title":"Foundry","type":"tags"},{"content":" Intro # \u0026ldquo;Vita, set the scene to microscopy.\u0026rdquo; The bench lights dim to a soft blue, the blinds close, and some calm music starts. Gloves stay on, nobody has to touch a button. Got your attention? 😉\nVoice is quickly becoming a normal way to talk to AI agents. The announcement post cites a study that found 14% of users already prefer speaking with generative AI over typing. I have been playing with voice for a while now. My first voice control experiments stitched together speech-to-text, an agent, and text-to-speech by hand. Later, in voice-agent, I compared Voice Live in \u0026ldquo;agent mode\u0026rdquo; with a native realtime model. Both worked, but I always had to assemble the voice part myself.\nHurray, that changes now! On September 24, 2026 Microsoft introduced native voice agents in Microsoft Foundry (currently in public preview). In this post I want to explain what a Foundry voice agent is and how it differs from the previous options. Then we build a friendly voice assistant for a simulated life science lab bench that controls lights, blinds, music, and a microscope camera. I kept the code as small as I could, and you can follow along with the companion repo foundry-voice-agent. Let\u0026rsquo;s dig in.\nOverview # What is a Foundry voice agent? # A voice agent is a new kind of agent in Foundry Agent Service, with kind: voice. Until now, a Foundry agent was a text agent, and voice was something you added around it. With a voice agent, voice is part of the agent definition itself. One versioned definition holds everything the service needs to run a spoken conversation:\nModel: a native speech-to-speech model such as gpt-realtime, or a supported text model that the service wraps with speech recognition and synthesis. The service picks the realtime or cascaded architecture based on the model you select. Model hosting: either a service-managed model (no deployment for you to look after) or your own self-deployed model. Behavior: instructions, plus an optional greeting that the agent speaks when a session starts. Listening: turn detection, noise reduction, echo cancellation, and transcription (including phrase lists for domain words). Speaking: the voice, its locale and speed, plus optional avatars. Tools and knowledge: the same tool and knowledge coverage as text agents (function tools, MCP, toolboxes, and system tools such as end_conversation). Data: whether conversations, transcripts, and audio are stored (store). Recording is off by default. Every create or update saves a new immutable version, and the agent\u0026rsquo;s endpoint is live as soon as the first version exists. There is no separate deployment step.\nOn top of that, the platform brings the things that usually take ages to build yourself. Voice agents have built-in telephony (inbound and outbound calls through Teams Phone extensibility and Twilio), voice-pipeline tracing that covers turn detection, audio input, model and tool calls, and audio output, and the same Foundry observability and evaluation experience as text agents. That includes rubric evaluators, which score each conversation against criteria you write for your agent.\nNOTE: you pick the interaction mode (text or voice) when you create an agent, and you can\u0026rsquo;t change it later. To give an existing text agent a voice, you create a new voice agent next to it.\nThree ways to build voice on Foundry # This is where it got a little confusing for me at first, because \u0026ldquo;Voice Live\u0026rdquo; and \u0026ldquo;voice agents\u0026rdquo; sound very similar. The announcement describes three approaches, and all three are still supported:\nVoice Live API (directly). You connect your app to the Voice Live API and configure everything in your session: model, instructions, voice, turn detection, and tools. Maximum flexibility, but your app owns the whole experience. This is the native-realtime example in my voice-agent repo. Voice Live with a Foundry text agent. Voice Live handles speech in and speech out around an existing Foundry text agent. The agent keeps its instructions, tools, and knowledge, and you tune the speech settings separately. This is the agent-mode example in my repo, and the topic of How to build a voice agent in the Speech docs. Foundry voice agents (new). Voice is native to the agent. Model, instructions, voice, greeting, turn-taking, and tools are managed together as one agent definition in Foundry. Under the hood it still uses Voice Live as the voice runtime. The announcement sums it up nicely: with the Voice Live API voice is a runtime, with Voice Live in front of a Foundry agent voice is a layer, and with a Foundry voice agent voice is the agent. In the diagram, watch where the voice part (blue) and the agent part (purple) live, and how much is left for your app (grey) to configure:\nflowchart TB subgraph A[\"1 · Voice is a runtime\"] direction LR A1[\"Your appmodel, instructions,voice, tools\"] \u003c--\u003e A2[\"Voice Live APIspeech in, model,speech out\"] end subgraph B[\"2 · Voice is a layer\"] direction LR B1[\"Your appvoice settings\"] \u003c--\u003e B2[\"Voice Livespeech to text,text to speech\"] \u003c--\u003e B3[\"Foundry text agentinstructions, tools,knowledge\"] end subgraph C[\"3 · Voice is the agent\"] direction LR C1[\"Your appjust audio\"] \u003c--\u003e C2 subgraph C2[\"Foundry voice agent (one definition)\"] direction TB C3[\"model, instructions, greeting,voice, turn-taking, tools, knowledge\"] C4[\"Voice Live runtime\"] end end A ~~~ B ~~~ C classDef app fill:#e7e5e4,stroke:#78716c,color:#1c1917 classDef voice fill:#bfdbfe,stroke:#3b82f6,color:#1c1917 classDef agent fill:#ddd6fe,stroke:#8b5cf6,color:#1c1917 class A1,B1,C1 app class A2,B2,C4 voice class B3,C3 agent style C2 fill:#ede9fe,stroke:#8b5cf6,color:#1c1917 The Microsoft Learn migration guide has a nice comparison between options 2 and 3. Here is my condensed version:\nVoice Live + Foundry text agent Foundry voice agent Models Text models of the connected agent Speech-to-speech models (e.g. gpt-realtime) and supported text models Model hosting The text agent\u0026rsquo;s model deployment Service-managed or self-deployed Conversation Speech to text, agent, text to speech Native speech-to-speech, or a text model with speech recognition and synthesis Voice configuration Voice Live settings, separate from the agent Model, instructions, voice, greeting, and turn-taking in one versioned definition Phone calls You integrate a telephony provider yourself Built-in Teams Phone extensibility and Twilio Conversation review Text history, which might differ from what the caller actually heard when they interrupted Opt-in transcripts, events, and audio Tracing The text-model interaction only The full voice pipeline, including turn detection and audio Python SDK azure-ai-voicelive azure-ai-projects[voice] (beta.voice_agents) So which one should you pick? My take (famous last words, \u0026ldquo;it depends\u0026rdquo;):\nStarting something new? Go with a Foundry voice agent. Microsoft recommends it as the path forward for new enterprise voice workloads, and you get telephony, voice tracing, and stored conversations out of the box instead of building them yourself. Already have a solid text agent? You don\u0026rsquo;t need to throw it away. A voice agent can call an existing text agent as a subagent for the specialist work, and the voice agent handles the conversation. Option 2 also keeps working. Need full control of every turn in your own code? Use the Voice Live API directly, or have a look at a hosted conversation engine behind a voice agent. Requirements # To follow along you need:\nA Microsoft Foundry project in a region where the voice agent preview is available The Foundry User role on that project Python 3.10 or later and the Azure CLI, signed in with az login Microsoft Edge or Google Chrome and a microphone (open speakers work, a headset is even better) The companion repo foundry-voice-agent Coding # The scenario: Vita, the lab bench assistant # In a life science lab you work with gloves on, often inside a biosafety cabinet. Touching a keyboard, a phone, or a light switch means taking gloves off or breaking your sterile workflow, which makes the lab bench a great fit for voice control. Our assistant Vita (Latin for \u0026ldquo;life\u0026rdquo;) can:\nTool What it does Try saying set_lights On or off, colour (white, warm, blue, green, red), dim 0 to 100% \u0026ldquo;Dim the lights to 30 percent\u0026rdquo; set_blinds Open, half, or closed \u0026ldquo;Close the blinds\u0026rdquo; set_scene Cell culture, microscopy, cleanup, end of day (lights, blinds, and music at once) \u0026ldquo;Set the scene to microscopy\u0026rdquo; play_music Play or pause the calm, focus, or upbeat playlist \u0026ldquo;Play some upbeat music\u0026rdquo; microscope_camera Start or stop a recording, take a snapshot \u0026ldquo;Take a snapshot\u0026rdquo; There is no real equipment involved. The lab is a web page (one SVG drawing) that redraws itself whenever a tool changes the state: the bench lights change colour and brightness, the room gets darker, the blinds move, an \u0026ldquo;IN PROGRESS\u0026rdquo; sign lights up during experiments, and the microscope camera monitor shows a blinking REC timer and snapshot thumbnails. For the music, the demo plays a small generative synth by default, so it works without any audio files. If you want real music, drop your own calm.mp3, focus.mp3, and upbeat.mp3 into static/music/ and they are picked up automatically. The repo README also explains why I didn\u0026rsquo;t go for Spotify or internet radio (spoiler: echo cancellation and licensing).\nNOTE: Vita\u0026rsquo;s instructions limit her to equipment commands, not experimental protocols or safety advice.\nIf your equipment exposes an API or an MCP server, the same pattern can connect voice commands to real actions, with appropriate authentication, permissions, and equipment safeguards. Other hands-free scenarios include adjusting room lighting in an operating room (subject to clinical safety and regulatory requirements), controlling a recording studio, or using an accessible smart home where reaching a switch isn\u0026rsquo;t practical.\nArchitecture # The app has three parts: the browser, a small FastAPI app, and the voice agent in Foundry.\nflowchart LR B[\"Browsermic, speaker, lab UI\"] \u003c-- \"audio + lab state(WebSocket)\" --\u003e R[\"app.pyFastAPI relay\"] R \u003c-- \"realtime session(Entra ID)\" --\u003e V[\"Foundry voice agentlab-voice-assistant\"] R --\u003e T[\"tools.pysimulated lab state\"] Why the relay in the middle? Two reasons. First, the voice agent endpoint needs a Microsoft Entra ID token, and the docs are clear: no bearer tokens in the URL and no long-lived credentials in browser code. The FastAPI app holds the credential and the browser only talks to localhost. Second, function tools run on the client side. The agent decides what to do, and our app actually does it and returns the result. As a bonus, the browser handles the microphone and speaker, so this also works fine in WSL without setting up any audio devices for Python.\nThe repo has only a handful of files:\nfoundry-voice-agent/ ├── instructions.md # Vita\u0026#39;s persona, written for speech ├── tools.py # simulated lab state + 5 tools + JSON schemas ├── create_agent.py # creates the voice agent in Foundry ├── app.py # FastAPI: UI + audio relay + tool calls └── static/ # lab UI (SVG), mic capture, playback └── music/ # optional: your own calm/focus/upbeat.mp3 1. A quick look in the portal # Before writing any code, it\u0026rsquo;s worth clicking through the portal once, because it shows nicely what a voice agent is made of. In the Foundry portal, open your project, select Build in the top navigation, go to Agents, and select New agent \u0026gt; Build an agent. In the Create an agent dialog, enter a name and choose Voice as the Interaction mode (marked as preview, and it can\u0026rsquo;t be changed later). In Voice agent goal you describe in plain words what the agent should do, and Foundry generates the instructions from it. Then select Create agent and open playground. Later, the Interaction type column (and filter) in the agent list shows you which of your agents are text and which are voice.\nThe playground shows the parts of the definition: Instructions (generated from the goal), a Greeting for the start of a session, the AI model (native speech or text model), the Voice, plus Avatar, Tools, Knowledge bases, and Advanced settings for input audio, transcription, and turn detection further down. In my project, the portal picked gpt-realtime-2.1 and the Ava Dragon HD Latest voice by default. Select Start, and you can talk to your agent right in the browser. Try interrupting it while it speaks to test barge-in. The tabs at the top (Traces, Monitor, Evaluation, and Channels for telephony) are where the operational side lives, and there\u0026rsquo;s even a Continue in code button.\nOpen the screenshot at full resolution to read the settings. TIP: the playground is perfect for tuning instructions and voices. For our lab assistant, though, I want the agent as code, so we can version it and recreate it at any time.\n2. The tools # tools.py is plain Python with no SDK in it. It holds the lab state, the five tool functions, and a JSON schema for each tool. Here is the schema for the scene tool and the dispatcher that app.py calls:\nSCENES = { \u0026#34;cell_culture\u0026#34;: {\u0026#34;lights\u0026#34;: {\u0026#34;power\u0026#34;: True, \u0026#34;color\u0026#34;: \u0026#34;white\u0026#34;, \u0026#34;brightness\u0026#34;: 100}, \u0026#34;blinds\u0026#34;: \u0026#34;half\u0026#34;, \u0026#34;music\u0026#34;: {\u0026#34;playing\u0026#34;: True, \u0026#34;playlist\u0026#34;: \u0026#34;focus\u0026#34;}}, \u0026#34;microscopy\u0026#34;: {\u0026#34;lights\u0026#34;: {\u0026#34;power\u0026#34;: True, \u0026#34;color\u0026#34;: \u0026#34;blue\u0026#34;, \u0026#34;brightness\u0026#34;: 15}, \u0026#34;blinds\u0026#34;: \u0026#34;closed\u0026#34;, \u0026#34;music\u0026#34;: {\u0026#34;playing\u0026#34;: True, \u0026#34;playlist\u0026#34;: \u0026#34;calm\u0026#34;}}, # ... cleanup, end_of_day } TOOL_SCHEMAS = [ # ... set_lights, set_blinds { \u0026#34;name\u0026#34;: \u0026#34;set_scene\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;Apply a lab scene that sets lights, blinds, and music together.\u0026#34;, \u0026#34;parameters\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;object\u0026#34;, \u0026#34;properties\u0026#34;: {\u0026#34;scene\u0026#34;: {\u0026#34;type\u0026#34;: \u0026#34;string\u0026#34;, \u0026#34;enum\u0026#34;: list(SCENES)}}, \u0026#34;required\u0026#34;: [\u0026#34;scene\u0026#34;], }, }, # ... play_music, microscope_camera ] def run_tool(name, args): \u0026#34;\u0026#34;\u0026#34;Execute a tool call from the voice agent and return a JSON-serializable result.\u0026#34;\u0026#34;\u0026#34; tool = _TOOLS.get(name) if tool is None: return {\u0026#34;ok\u0026#34;: False, \u0026#34;message\u0026#34;: f\u0026#34;Unknown tool \u0026#39;{name}\u0026#39;.\u0026#34;} # ... validate enum values, then call the function return tool(**args) Every tool returns a small result like {\u0026quot;ok\u0026quot;: true, ...}. That matters, because the agent\u0026rsquo;s instructions tell it to only confirm a change after the tool returned ok.\nThe instructions in instructions.md are written for speech, not for a chat window: one short sentence per reply, no lists, no markdown, act right away on clear requests, and ask for a quick confirmation before stopping a microscope recording. It also says what Vita must never do (give advice on protocols, safety procedures, or hazardous materials) and what to say when a tool fails.\n3. Create the voice agent as code # Install the dependencies first. The voice extra of the Foundry SDK adds the WebSocket libraries for the realtime connection:\ngit clone https://github.com/beyondelastic/foundry-voice-agent.git cd foundry-voice-agent python -m venv .venv \u0026amp;\u0026amp; source .venv/bin/activate pip install -r requirements.txt # azure-ai-projects[voice]\u0026gt;=2.7.0, azure-identity, fastapi, ... cp .env.example .env # set FOUNDRY_PROJECT_ENDPOINT az login create_agent.py builds a VoiceAgentDefinition. If you have created a prompt agent with the SDK before, this will look very familiar. The difference is all the voice-specific parts:\ndefinition = VoiceAgentDefinition( model_type=VoiceModelType.MANAGED, model=model, # gpt-realtime by default instructions=(Path(__file__).parent / \u0026#34;instructions.md\u0026#34;).read_text(encoding=\u0026#34;utf-8\u0026#34;), greeting=VoiceAgentTemplateGreetingConfig( text=\u0026#34;Hi, I\u0026#39;m Vita, your lab bench assistant. What can I set up for you?\u0026#34; ), audio=VoiceAgentAudioConfig( input=VoiceAgentAudioInputConfig(...), # echo cancellation + turn detection, see step 6 output=VoiceAgentAudioOutputConfig(voice=voice, voice_type=VoiceType.AZURE_STANDARD), ), output_modalities=[VoiceOutputModality.AUDIO], # Function tools are executed by our app (app.py), not by the service. tools=[ VoiceAgentFunctionTool( name=tool[\u0026#34;name\u0026#34;], description=tool[\u0026#34;description\u0026#34;], parameters=RealtimeFunctionToolParameters(tool[\u0026#34;parameters\u0026#34;]), ) for tool in TOOL_SCHEMAS ], store=True, ) with ( DefaultAzureCredential() as credential, AIProjectClient(endpoint=endpoint, credential=credential, allow_preview=True) as project_client, ): created = project_client.agents.create_version(agent_name=agent_name, definition=definition) A few things worth pointing out:\nmodel_type=VoiceModelType.MANAGED uses a service-managed model. There is no model deployment to create. To use your own deployment, switch to VoiceModelType.SELF_DEPLOYED and pass your deployment name. The greeting is a fixed template here, so Vita always opens with the same line. You can also let the model write it with VoiceAgentLlmGeneratedGreetingConfig. store=True keeps the conversation (transcript and audio) so you can review it later. audio.input controls how the agent listens. We\u0026rsquo;ll get back to it in step 6, because it is the key to stopping Vita from interrupting herself. allow_preview=True is required because voice agents are a preview feature. Run it, and the agent shows up in your project with its tools attached:\npython create_agent.py # Voice agent \u0026#39;lab-voice-assistant\u0026#39; saved as version 1 (model: gpt-realtime, voice: en-US-AvaNeural) NOTE: I tested the demo with the managed gpt-realtime model. The portal picked gpt-realtime-2.1 by default (see step 1), so check the Models page of your project for what\u0026rsquo;s available in your region and set FOUNDRY_VOICE_AGENT_MODEL in .env.\n4. The relay and the tool loop # app.py is where things get interesting. When the browser connects to /ws, the app opens a realtime session with the voice agent:\nasync with client.beta.voice_agents.realtime.connect(agent_name=AGENT_NAME) as agent: ... That is the whole connection setup. There is no session.update with model, voice, or instructions, because the agent already owns its configuration in Foundry. Under the hood the SDK connects to the project endpoint (.../api/projects/\u0026lt;project\u0026gt;/agents/\u0026lt;name\u0026gt;/endpoint/protocols/voice) with an Entra ID token and adds the preview opt-in header for you.\nTwo small tasks then run side by side. One forwards microphone audio from the browser to the agent:\nasync def browser_to_agent(browser: WebSocket, agent): while True: audio = await browser.receive_bytes() await agent.input_audio_buffer.append(audio=audio) The other loops over the agent\u0026rsquo;s events. Most of them are simply passed to the browser: audio chunks, transcripts, and a \u0026ldquo;the user started talking\u0026rdquo; signal for barge-in. The interesting part is the function call:\nasync for event in agent: if isinstance(event, RealtimeServerEventResponseAudioDelta): await browser.send_bytes(event.delta) # PCM16, mono, 24 kHz elif isinstance(event, RealtimeServerEventInputAudioBufferSpeechStarted): # Barge-in: the user started talking, so stop the agent\u0026#39;s current reply. await browser.send_json({\u0026#34;type\u0026#34;: \u0026#34;interrupt\u0026#34;}) if response_active: await agent.response.cancel() elif isinstance(event, RealtimeServerEventResponseFunctionCallArgumentsDone): args = json.loads(event.arguments or \u0026#34;{}\u0026#34;) result = tools.run_tool(event.name, args) pending_outputs.append((event.call_id, result)) await browser.send_json({\u0026#34;type\u0026#34;: \u0026#34;state\u0026#34;, \u0026#34;state\u0026#34;: tools.STATE}) elif isinstance(event, RealtimeServerEventResponseDone): response_active = False if pending_outputs: for call_id, result in pending_outputs: await agent.conversation.item.create( item=RealtimeConversationItemFunctionCallOutput(call_id=call_id, output=json.dumps(result)) ) pending_outputs.clear() await agent.response.create() # ... transcripts, response.created, errors Here is what happens when you say \u0026ldquo;set the scene to microscopy\u0026rdquo;:\n%%{init: {\"themeVariables\": {\"signalColor\": \"#10b981\", \"signalTextColor\": \"#10b981\", \"actorLineColor\": \"#a8a29e\", \"sequenceNumberColor\": \"#1c1917\"}, \"sequence\": {\"messageFontWeight\": 600}}}%% sequenceDiagram participant B as Browser participant A as app.py participant V as Voice agent B-\u003e\u003eA: mic audio A-\u003e\u003eV: input_audio_buffer.append V-\u003e\u003eA: function_call_arguments.done (set_scene) A-\u003e\u003eA: run_tool() updates lab state A-\u003e\u003eB: new state (lab redraws) V-\u003e\u003eA: response.done A-\u003e\u003eV: function_call_output + response.create V-\u003e\u003eA: audio \"Microscopy scene is set.\" A-\u003e\u003eB: audio TIP: send the tool output only after the response.done of the function-call response. If you call response.create() while that response is still finishing, you can run into a \u0026ldquo;concurrent response\u0026rdquo; error. The official function tool sample uses the same pattern.\n5. The browser side # The browser does three jobs, all in plain JavaScript in static/:\nCapture: a tiny AudioWorklet turns the microphone stream into 100 ms chunks of PCM16 at 24 kHz, the format the agent expects, and sends them over the WebSocket. When client-reference echo cancellation is enabled, it adds a second channel containing the audio sent to the speakers; more on that in step 6. Playback: audio chunks from the agent are queued back to back with the Web Audio API. When the app detects that you’ve started speaking, it sends an interrupt message and the browser stops queued playback. Rendering: every state message updates a few CSS variables (light colour and brightness, darkness, blinds position) that the SVG lab uses, plus the status bar, the music, and the microscope camera widgets. A small orb next to Vita\u0026rsquo;s name pulses with your voice (teal) or hers (purple). 6. Don\u0026rsquo;t let Vita interrupt herself # When I first tested the demo with open laptop speakers, the assistant sometimes stopped in the middle of a sentence. The reason: the microphone picked up her own voice (and the music), the turn detection thought I was talking, and the agent did exactly what it should do on a barge-in. It stopped talking. 🙃\nMy first thought was: \u0026ldquo;The browser has echo cancellation, right?\u0026rdquo; It does, but with echoCancellation: true a browser must attempt to cancel at least audio from remote WebRTC tracks, and should attempt to cancel all system audio (MDN). That doesn\u0026rsquo;t guarantee it will cancel audio this page plays through the Web Audio API (exactly how we play Vita\u0026rsquo;s voice). In my tests, the browser\u0026rsquo;s echo cancellation wasn\u0026rsquo;t enough.\nThe browser plays Vita\u0026rsquo;s voice and the music through the same audio output, while the microphone can pick up both. We use live-reference echo cancellation (Voice Live calls it Live-Reference AEC) so the service can distinguish speaker audio from the caller\u0026rsquo;s voice. The browser sends its microphone audio on channel 0 and the audio it plays on channel 1; the voice agent uses channel 1 as a reference to reduce echo in channel 0. On the agent side, we combine it with noise suppression and semantic turn detection configured to remove supported filler words, which can help reduce false barge-ins:\ninput=VoiceAgentAudioInputConfig( format=RealtimeAudioFormatsAudioPcm(rate=24000), # channel 0 = microphone, channel 1 = what the speakers are playing echo_cancellation=VoiceAgentEchoCancellation(reference_source=\u0026#34;client\u0026#34;, channels=2), noise_reduction=VoiceAgentNoiseReduction(type=\u0026#34;azure_deep_noise_suppression\u0026#34;), turn_detection=VoiceAgentAzureSemanticVadTurnDetection( threshold=0.6, speech_duration_ms=timedelta(milliseconds=200), silence_duration_ms=timedelta(milliseconds=500), remove_filler_words=True, ), ), In the browser, everything we play (Vita and the music) goes through one speakers node. That node feeds the real speakers and the second input of the capture worklet:\nspeakers.connect(ctx.destination); // what you hear micSource.connect(capture, 0, 0); // worklet input 0: microphone speakers.connect(capture, 0, 1); // worklet input 1: echo reference The worklet then interleaves both into one PCM16 stream:\n// mic-worklet.js: [mic, speakers] per frame for (let i = 0; i \u0026lt; mic.length; i++) { this.buffer[this.length++] = pcm16(mic[i]); if (this.channels === 2) this.buffer[this.length++] = speakers ? pcm16(speakers[i]) : 0; if (this.length === this.buffer.length) { this.port.postMessage(this.buffer.buffer.slice(0)); this.length = 0; } } Mono versus stereo has to match the agent definition, otherwise the service hears garbage. So at the start of each session app.py reads the agent\u0026rsquo;s latest version and tells the browser how many channels to send. As a small bonus, the music also drops to a low volume (ducking) while Vita talks.\nTIP: if Vita still gets interrupted in a loud room, raise the threshold a bit. This demo\u0026rsquo;s app.py also explicitly cancels an active reply when it receives the speech-start event, so setting interrupt_response=False alone won\u0026rsquo;t turn off barge-in; you\u0026rsquo;d also need to disable that app-side cancellation. A headset is still the most reliable fix of all.\n7. Run it # uvicorn app:app --reload Open http://localhost:8000, select Start session, allow the microphone, and Vita greets you. Now try a few things:\n\u0026ldquo;Set the scene to microscopy.\u0026rdquo; \u0026ldquo;Make the lights blue and dim them to 30 percent.\u0026rdquo; \u0026ldquo;Start recording.\u0026rdquo; \u0026hellip; \u0026ldquo;Take a snapshot.\u0026rdquo; \u0026hellip; \u0026ldquo;Stop the recording.\u0026rdquo; (Vita asks for a quick confirmation first) \u0026ldquo;Play some upbeat music.\u0026rdquo; and then interrupt Vita while she is answering. Here is Vita in action. Turn on your sound to hear the conversation:\nYour browser doesn't support embedded video. Download the demo recording. NOTE: the demo listens all the time and reacts to everything it hears. In a real lab, where people talk across benches all day, I\u0026rsquo;d add a wake word as a gate: the app only streams microphone audio to the voice agent after someone says \u0026ldquo;Hey Vita\u0026rdquo;. Azure AI Speech custom keyword would have been my first pick, but Speech Studio currently shows a notice that its model training will be retired on August 1, 2027 (existing models keep working). For a new project I\u0026rsquo;d rather look at Picovoice Porcupine, whose web SDK runs on-device in the browser (custom wake words via the Picovoice Console, AccessKey required), or the open-source openWakeWord for Python, which lets you train your own wake word from synthetic speech.\nTIP: whenever you change instructions.md, the tools, or the audio settings, just run python create_agent.py again. It saves a new agent version, and the next session uses it. No need to restart the app.\nObservability and evaluation # Because we set store=True, every session is saved as a conversation that you can read back by ID with project_client.beta.voice_agents.conversations. Voice agents also use the same Foundry observability experience as other agents, so each conversation shows up as a trace that covers the voice pipeline too. And this is where it gets really useful for a scenario like ours: with rubric evaluators you can score conversations against your own criteria, for example \u0026ldquo;asked for confirmation before stopping a recording\u0026rdquo; or \u0026ldquo;never gave safety or protocol advice\u0026rdquo;. Foundry can even generate the rubric from the agent\u0026rsquo;s instructions and production traces. I\u0026rsquo;ll explore that in a future post.\nClosing # Voice agents are a big step for Foundry. Until now, adding voice meant choosing between flexibility (Voice Live directly) and reusing a governed agent (Voice Live with a text agent), and either way you assembled and observed the voice part yourself. Now voice is a first-class agent type: one versioned definition for model, voice, greeting, turn-taking, and tools, plus telephony, tracing, and evaluation built in. For our lab assistant that meant very little code: a definition, five plain Python functions, and a small relay loop. Even the trickiest part, keeping Vita from hearing herself through open speakers, was mostly configuration plus a second audio channel from the browser.\nKeep in mind that voice agents are in public preview, so APIs and defaults can still change, and you should test quality, latency, and cost for your own workload. From here, there are plenty of next steps to try: swap the simulated tools for real device APIs through an MCP server or a toolbox, add a phrase list for lab vocabulary (reagents, cell lines, instruments), give Vita an avatar, or let her answer a phone line. If you build something with it, let me know. Happy talking!\nSources # foundry-voice-agent (companion repo) Introducing voice agents in Microsoft Foundry (Microsoft Community Hub) Quickstart: Create a voice-based prompt agent (Microsoft Learn) Configure a voice agent (Microsoft Learn) Migrate from Voice Live with Foundry Agent Service to Microsoft Foundry voice agents (Microsoft Learn) How to build a voice agent with Voice Live and Foundry Agent Service (Microsoft Learn) Voice Live API overview (Microsoft Learn) How to use the Voice Live API, including Live-Reference AEC (Microsoft Learn) Keyword recognition overview (Microsoft Learn) Picovoice Porcupine (GitHub) and openWakeWord (GitHub) MediaTrackConstraints: echoCancellation (MDN) Function tool sample for voice agents (GitHub) Live audio conversation sample (Azure SDK for Python) voice-agent (my earlier Voice Live experiments) ","date":"9 October 2026","externalUrl":null,"permalink":"/posts/foundry-voice-agents/","section":"Posts","summary":"Intro # \u0026ldquo;Vita, set the scene to microscopy.","title":"Native Voice Agents in Microsoft Foundry: When Speech and Agent Become One","type":"posts"},{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/posts/","section":"Posts","summary":"","title":"Posts","type":"posts"},{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/tags/","section":"Tags","summary":"","title":"Tags","type":"tags"},{"content":"","date":"9 October 2026","externalUrl":null,"permalink":"/tags/voice/","section":"Tags","summary":"","title":"Voice","type":"tags"},{"content":" Intro # Here is a line I keep coming back to: the agent is code, not clicks. It is easy to build an agent by dragging things around in a portal. It is a lot harder to answer the boring production questions that follow: which version is running, who reviewed the last prompt change, and how do I roll it back? The moment an agent goes to production, it has to live where the rest of your system lives, in Git, behind a pipeline.\nThat is exactly what declarative agents in Microsoft Foundry give you. In this post I want to unpack what \u0026ldquo;declarative\u0026rdquo; means for both prompt agents and hosted agents, why treating your agent as code is such a big deal, and how to continuously deploy one with GitHub Actions, including the app registration, OIDC, and role setup the pipeline actually needs. To keep it copy-and-runnable I built a tiny companion repo, foundry-declarative-agent, that you can clone and follow along with. The example agent, TrialFinder, uses Foundry\u0026rsquo;s native web search tool to surface currently recruiting clinical trials from public sources. Let\u0026rsquo;s dig in.\nOverview # What makes an agent \u0026ldquo;declarative\u0026rdquo;? # Foundry gives you two shapes of agent, and both can be declarative:\nPrompt agents are defined entirely by configuration: a model deployment, a set of instructions, and some tools. There is no application code to run, Foundry runs the agent for you. The whole definition is a small artifact you can put in a file. Hosted agents are your own code (Python or C#) that Foundry hosts and scales. They are still declared through a committed azure.yaml and deployed with a pipeline. I covered these in depth in my hosted agents post. The unifying idea is the important part: the agent\u0026rsquo;s definition is a versioned artifact, not a state you poke into a UI. In the companion repo, that artifact is a single file, .foundry/agent-metadata.yaml:\napiVersion: foundry/v1 kind: PromptAgent agentName: trial-finder description: Finds currently recruiting clinical trials from public sources via web search. # Instructions live in their own reviewable file instructions_file: prompts/system.md # Native, out-of-the-box web search tool (no connection to provision) tools: - type: web_search That is the entire agent: its instructions, a name, and a tool. The web search tool is native to Foundry, no application to host and no connection to wire up, so the agent gains a real capability without any infrastructure. The model deployment is supplied at deploy time from an environment variable, and the instructions sit in their own prompts/system.md, so a prompt tweak is a one-line diff you can review like any other change.\nNOTE: the native web search tool (\u0026ldquo;Search the web with Bing Search\u0026rdquo; in the portal) needs no setup. Only the domain-restricted variant, Bing Custom Search, requires a connection. Either way it is billed through Grounding with Bing, so it is not free.\nWhy bother? The benefits of agents as code # Once the agent is a file in your repo, it inherits everything Git and CI already give the rest of your app:\nReproducible. Every deployed agent maps back to a Git SHA. \u0026ldquo;What\u0026rsquo;s in production?\u0026rdquo; has a real answer. Reviewable. A prompt change, a new tool, a model bump, each one is a pull request with a diff and an approver, not a silent portal edit. Rollback-friendly. Foundry keeps immutable agent versions, so reverting is a revert, not a reconstruction from memory. No portal drift. The running agent can\u0026rsquo;t quietly diverge from what\u0026rsquo;s in source control, because source control is what deploys it. Same pipeline as the app. The agent gets the exact CI/CD treatment as the rest of your app: pull request, review, merge, deploy. Auditable and promotable. Point the same definition at a different project or model via config, so promoting an agent across dev and prod is a variable change, not a manual rebuild. Agents as microservices # Here is the framing I find most useful: treat the agent like a microservice. It is not code you embed in your app, it is a separate, independently deployed unit your app calls:\nIt has its own contract, a versioned definition (this YAML) and a stable endpoint other services invoke. It has its own identity: it runs under a Microsoft Entra agent identity (your project\u0026rsquo;s shared agent identity by default, or a dedicated one once you publish it), and, in CI, the pipeline\u0026rsquo;s own identity at deploy time. It versions independently, editing the prompt ships a new agent version without touching or redeploying the calling app. It ships through its own pipeline, the same PR-and-merge flow as any other service. Independently deployable, contract-first, own identity, own lifecycle. That is a microservice, it just happens to reason with a model.\nRequirements # To follow along with the companion repo:\nA Microsoft Foundry project with a chat model deployment (I use gpt-5.4-mini). The Foundry User role on your Foundry resource for your own user principal, this is what grants the data-plane agents/write action that creating an agent version needs. If you created the project as an Azure Owner (and so can assign roles), Foundry adds this for you automatically; a plain Contributor, or the pipeline\u0026rsquo;s service principal later, won\u0026rsquo;t have it, so assign it explicitly. The web search tool available on your project. General web search is native (nothing to provision), but it is billed via Grounding with Bing and an admin can enable or disable it. The Azure CLI and the GitHub CLI (handy for the CI setup below). Python 3.12+. A GitHub repo, we\u0026rsquo;ll wire it to Azure with OIDC / workload identity federation, no long-lived secrets. To run the CI setup below, permission to register an Entra app (the Application Developer role, or a tenant that allows app registrations), to assign Azure roles on the Foundry resource (Owner, User Access Administrator, or Role Based Access Control Administrator), and admin on the GitHub repo (to set Actions secrets and variables). Coding # 1. The agent lives in the repo # We already saw .foundry/agent-metadata.yaml above. The key move is that this file, plus prompts/system.md, is the agent. Nobody authors it in a UI, and honestly I don\u0026rsquo;t want my agents authored in a UI anyway, I want them in Git where every change is reviewable.\n2. Create a version locally # Before wiring any pipeline, prove it locally. A small script, scripts/sync_agent.py, turns the YAML into an agent version. Condensed to its essentials, it reads the declarative file, builds a PromptAgentDefinition, and calls create_version:\n# --- Illustrative excerpt only, run scripts/sync_agent.py in the repo --- from azure.ai.projects import AIProjectClient from azure.ai.projects.models import PromptAgentDefinition, WebSearchTool from azure.identity import DefaultAzureCredential meta = yaml.safe_load(META_PATH.read_text()) # the declarative agent-metadata.yaml definition = PromptAgentDefinition( model=os.environ[\u0026#34;AZURE_AI_MODEL_DEPLOYMENT\u0026#34;], instructions=instructions, # loaded from prompts/system.md tools=[WebSearchTool()], # the `- type: web_search` entry ) client = AIProjectClient(endpoint=endpoint, credential=DefaultAzureCredential()) version = client.agents.create_version(agent_name=agent_name, definition=definition) # --- End illustrative excerpt --- Three things worth noting: the tools: list in the YAML maps to SDK tool objects (web_search becomes WebSearchTool()), DefaultAzureCredential means the same code authenticates with your az login locally and with the pipeline\u0026rsquo;s identity in CI, and every create_version call produces a new immutable version. The excerpt above trims the env-var checks and the YAML-to-tools mapping for readability, so run the full sync_agent.py, not this snippet, but the shape is exactly that. It reads its config from the environment, so on your machine it just uses your az login:\npython -m venv .venv \u0026amp;\u0026amp; source .venv/bin/activate pip install -r requirements.txt cp .env.example .env # set your project endpoint + model deployment set -a \u0026amp;\u0026amp; source .env \u0026amp;\u0026amp; set +a az login python scripts/sync_agent.py # -\u0026gt; Created agent version: name=trial-finder version=1 Open the agent in the Foundry playground and ask it something like \u0026ldquo;Find currently recruiting phase 3 melanoma immunotherapy trials in Europe\u0026rdquo;, it web-searches and answers.\nEdit prompts/system.md, run the script again, and you get version 2. That is the whole thesis: a prompt diff produces a new immutable version, no portal.\nNOTE: if you hit a 403 like \u0026ldquo;does not have permissions for \u0026hellip;agents/write\u0026rdquo;, that is the data-plane role. Assign yourself Foundry User on the Foundry resource and retry.\n3. Set up continuous deployment (GitHub Actions) # The pipeline runs the same sync_agent.py, but a fresh GitHub runner has no az login. So we give the pipeline its own identity and let it authenticate with OIDC, no stored secret. It is four one-time steps.\na. Create an app registration. The identity CI acts as:\nAPP_ID=$(az ad app create --display-name \u0026#34;foundry-declarative-agent-ci\u0026#34; --query appId -o tsv) az ad sp create --id \u0026#34;$APP_ID\u0026#34; b. Add a federated credential. So Entra trusts GitHub\u0026rsquo;s token instead of a secret. The one field that must match exactly is the subject, and GitHub will tell you what it emits, so read it rather than guess it:\n# GitHub\u0026#39;s exact subject prefix for your repo (modern accounts embed numeric IDs) SUBJECT=\u0026#34;$(gh api /repos/beyondelastic/foundry-declarative-agent/actions/oidc/customization/sub --jq .sub_claim_prefix):ref:refs/heads/main\u0026#34; echo \u0026#34;$SUBJECT\u0026#34; # e.g. repo:beyondelastic@11375951/foundry-declarative-agent@1323781534:ref:refs/heads/main az ad app federated-credential create --id \u0026#34;$APP_ID\u0026#34; --parameters \u0026#34;$(jq -n --arg sub \u0026#34;$SUBJECT\u0026#34; \u0026#39;{ name: \u0026#34;github-main\u0026#34;, issuer: \u0026#34;https://token.actions.githubusercontent.com\u0026#34;, subject: $sub, audiences: [\u0026#34;api://AzureADTokenExchange\u0026#34;] }\u0026#39;)\u0026#34; Replace beyondelastic/foundry-declarative-agent with your own \u0026lt;owner\u0026gt;/\u0026lt;repo\u0026gt;. Reading the subject from sub_claim_prefix is the reliable move because repositories created after July 15, 2026 use GitHub\u0026rsquo;s immutable subject format with owner and repo IDs (repo:owner@ownerID/repo@repoID:ref:...), while older repos use the legacy repo:owner/repo:..., sub_claim_prefix gives you the right one either way. Merging a PR into main counts as a push to main, so this one credential covers the whole PR-and-merge flow. You\u0026rsquo;d only add a :pull_request credential (and a pull_request trigger) if you also run the workflow on the PR itself before merge.\nNOTE: if azure/login fails later with AADSTS700213: No matching federated identity record, the subject didn\u0026rsquo;t match. The run log prints the exact subject claim just above the error, copy that literal value into the credential\u0026rsquo;s subject.\nc. Assign the role. The same Foundry User you gave yourself, now for the pipeline\u0026rsquo;s identity:\nSP_OID=$(az ad sp show --id \u0026#34;$APP_ID\u0026#34; --query id -o tsv) az role assignment create \\ --assignee-object-id \u0026#34;$SP_OID\u0026#34; \\ --assignee-principal-type ServicePrincipal \\ --role \u0026#34;Foundry User\u0026#34; \\ --scope \u0026#34;/subscriptions/\u0026lt;sub\u0026gt;/resourceGroups/\u0026lt;rg\u0026gt;/providers/Microsoft.CognitiveServices/accounts/\u0026lt;foundry-account\u0026gt;\u0026#34; d. Tell GitHub who to be and what to target. The three IDs go in secrets (they are identifiers, not passwords, the OIDC exchange does the real auth); the config goes in variables:\ngh secret set AZURE_CLIENT_ID --body \u0026#34;$APP_ID\u0026#34; gh secret set AZURE_TENANT_ID --body \u0026#34;$(az account show --query tenantId -o tsv)\u0026#34; gh secret set AZURE_SUBSCRIPTION_ID --body \u0026#34;$(az account show --query id -o tsv)\u0026#34; gh variable set AZURE_AI_PROJECT_ENDPOINT --body \u0026#34;https://\u0026lt;res\u0026gt;.services.ai.azure.com/api/projects/\u0026lt;project\u0026gt;\u0026#34; gh variable set AZURE_AI_MODEL_DEPLOYMENT --body \u0026#34;gpt-5.4-mini\u0026#34; gh variable set AZURE_FOUNDRY_AGENT_NAME --body \u0026#34;trial-finder\u0026#34; With that in place, .github/workflows/deploy.yml runs the sync on every merge:\nname: deploy-agent on: push: branches: [main] paths: [\u0026#39;.foundry/**\u0026#39;, \u0026#39;scripts/sync_agent.py\u0026#39;] workflow_dispatch: permissions: id-token: write # lets the job request the OIDC token contents: read jobs: sync-agent: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: python-version: \u0026#34;3.12\u0026#34; - run: pip install -r requirements.txt - name: Azure login (OIDC) uses: azure/login@v2 with: client-id: ${{ secrets.AZURE_CLIENT_ID }} tenant-id: ${{ secrets.AZURE_TENANT_ID }} subscription-id: ${{ secrets.AZURE_SUBSCRIPTION_ID }} - name: Reconcile the declarative agent env: AZURE_AI_PROJECT_ENDPOINT: ${{ vars.AZURE_AI_PROJECT_ENDPOINT }} AZURE_AI_MODEL_DEPLOYMENT: ${{ vars.AZURE_AI_MODEL_DEPLOYMENT }} AZURE_FOUNDRY_AGENT_NAME: ${{ vars.AZURE_FOUNDRY_AGENT_NAME }} run: python scripts/sync_agent.py The three moving parts: permissions: id-token: write lets the job mint the OIDC token, azure/login@v2 exchanges it for an Azure token as your app, and DefaultAzureCredential in the script picks that up, exactly like your laptop run, minus the secret.\nflowchart LR Dev[Developer] --\u003e|PR merge to main| Repo[GitHub repo] Repo --\u003e|deploy.yml| GA[GitHub ActionsOIDC then azure/login] GA --\u003e|python sync_agent.py| Foundry[Foundrynew agent version] 4. The payoff: change the prompt, ship a version # This is where it clicks. Say compliance decides the informational disclaimer should appear on every reply, not just the ones that name a specific trial. In .foundry/prompts/system.md that is a one-line change:\n- End any response that names specific trials with: *Informational only ...* + End every response with: *Informational only ...* Open a PR, get it reviewed, and merge. The push to main triggers the workflow, which runs the same sync_agent.py, and Foundry gets a new agent version, all from a Git diff, no portal clicks.\nAsk the agent anything in the playground now and the disclaimer is there every time. A product ask, delivered as a reviewed change with an audit trail instead of a portal edit nobody can trace.\nTIP: your local run and the CI job call the exact same sync_agent.py, so \u0026ldquo;works on my machine\u0026rdquo; and \u0026ldquo;works in the pipeline\u0026rdquo; are, for once, the same code path.\nClosing # Declarative agents are the natural endpoint of taking agents seriously as production software. Once the definition is a file, everything good about modern DevOps just applies: pull requests, reproducible deploys, rollbacks, environment promotion, and the same pipeline for your agent as for your API. Prompt agent or hosted agent, the agent stops being a special snowflake in a portal and becomes what it always should have been, a versioned, reviewable, independently deployable service.\nIf you want to build it yourself, clone the minimal foundry-declarative-agent and run through it top to bottom, it is the exact flow from this post. And if you want the other half of the story, bringing your own code as a hosted agent, that\u0026rsquo;s my hosted agents post. Happy shipping!\nSources # foundry-declarative-agent (companion repo) Create a prompt agent (Microsoft Learn quickstart) Web search tool in Foundry Agent Service (Microsoft Learn) Role-based access control for Microsoft Foundry (Microsoft Learn) Hosted agents in Foundry Agent Service (Microsoft Learn) Workload identity federation (Microsoft Learn) Security hardening with OpenID Connect (GitHub Docs) Azure CLI GitHub CLI ","date":"3 August 2026","externalUrl":null,"permalink":"/posts/declarative-agents-foundry/","section":"Posts","summary":"Intro # Here is a line I keep coming back to: the agent is code, not clicks.","title":"Declarative Agents in Microsoft Foundry: Agents as Code, Deployed by GitHub Actions","type":"posts"},{"content":"","date":"3 August 2026","externalUrl":null,"permalink":"/categories/devops/","section":"Categories","summary":"","title":"DevOps","type":"categories"},{"content":"","date":"3 August 2026","externalUrl":null,"permalink":"/tags/github/","section":"Tags","summary":"","title":"GitHub","type":"tags"},{"content":"","date":"22 July 2026","externalUrl":null,"permalink":"/tags/agent-framework/","section":"Tags","summary":"","title":"Agent Framework","type":"tags"},{"content":" Intro # Building an agent on your laptop is the easy part. Getting that same agent to run securely for thousands of users, with its own identity, persistent state, and proper observability? That is where most projects hit a wall. I have spent a lot of time lately building agents, and the pattern is always the same: the intelligence is done in an afternoon, and then the \u0026ldquo;make it production-ready\u0026rdquo; work eats the next two weeks.\nHosted agents in Microsoft Foundry Agent Service are Microsoft\u0026rsquo;s answer to exactly that gap. In this post I want to walk through what hosted agents are, why I think they are a big deal, and then get hands-on with three real examples, from your first hosted agent to bringing your own framework. Let\u0026rsquo;s dig in.\nOverview # A hosted agent is a containerized agentic application that Foundry runs and manages for you. You write the agent logic (in Python or C#), pick whatever framework you like, and deploy it. Foundry provisions the compute, gives the agent a dedicated identity, exposes a stable endpoint, and handles scaling, state, and tracing. You bring the code, the platform handles the plumbing.\nThe important word here is \u0026ldquo;your code\u0026rdquo;. Unlike prompt-based agents that you define entirely through prompts and tool configuration in the portal, a hosted agent is your own application. You choose the framework (Microsoft Agent Framework, LangGraph, Semantic Kernel, or plain custom code), you control the runtime behavior, and you deploy a container image (or, more recently, just your source code) to Microsoft-managed infrastructure.\nMicrosoft announced the public preview refresh of hosted agents earlier this year and followed up with a batch of updates around Microsoft Build. The feature moved quickly, and hosted agents reached general availability at the beginning of July 2026. Several of the newer capabilities on top (like source-code deployment and the agent optimizer) are still in public preview, which I will flag as we go.\nWhy not just use a container app or a function? # This was my first question too. Traditional compute (web apps, containers, serverless functions) was built for web services where many users share the same instance. That is perfectly fine for a REST API. It is a problem for agents.\nThink about what an agent harness actually does: it reads and writes files, it executes arbitrary code, it holds sensitive context and credentials. Now imagine Customer A and Customer B both hitting the same shared container. Suddenly one session can see another session\u0026rsquo;s files and state. That is not just inefficient, it is a security and isolation nightmare.\nHosted agents solve this with per-session, VM-isolated sandboxes. Every single agent session gets its own hypervisor-isolated sandbox with a persistent filesystem. Here is how the two approaches compare:\nTraditional compute Hosted agents Isolation Many sessions share a container Every session gets a dedicated sandbox Cold starts Seconds to minutes, high variance Seconds, low variance, predictable Idle cost Always-on billing or slow scale-from-zero Scale to zero with filesystem-preserving resume State You build it (databases, external storage) Built in, files and disk survive scale-to-zero Identity Shared service account Per-agent Microsoft Entra ID (agent identity) Observability You build it Built in (agent, session, fleet) How it works # You package your agent as a container image and push it to Azure Container Registry (or deploy straight from source, more on that later). When you deploy, Agent Service pulls the image, provisions compute, assigns a dedicated Microsoft Entra ID, and exposes a dedicated endpoint. At runtime your code handles requests and can call Foundry models, Toolbox tools, and downstream Azure services using its agent identity.\nflowchart LR Dev[\"Developer (local)\"] --\u003e|azd deploy / provision| Reg[\"Container registry(managed on code, your ACR on container)\"] Reg --\u003e|image pull| Host[\"Foundry hosted runtime\"] Host --\u003e|Responses / Invocations| Client[\"Portal / SDK / Teams\"] Host --\u003e|agent identity| Model[\"Foundry model\"] Host --\u003e|agent identity| Tools[\"Toolbox / Azure services\"] Real-world examples # This clicks faster with concrete scenarios. A few that map neatly onto hosted agents:\nA coding agent that refactors a repo overnight. It clones the code into its sandbox, writes and executes code, and needs that working directory to survive across a long-running task. Per-session filesystem persistence is exactly what it wants. A customer support agent published to Teams. It answers thousands of conversations in parallel, each one isolated, each one carrying the user\u0026rsquo;s identity through On-Behalf-Of so it only sees what that user is allowed to see. A research agent that synthesizes hundreds of documents into a briefing. It downloads files into $HOME, processes them, and can scale to zero between runs so you pay nothing while it waits for the next request. A webhook receiver for GitHub, Stripe, or Jira. These systems send their own payload format, so the agent uses the Invocations protocol to accept arbitrary JSON rather than an OpenAI-style chat message. Key concepts to know # Before we write code, a few concepts that shape how you build:\nSessions and conversations. A session is a logical unit with persisted state ($HOME plus files uploaded via the /files endpoint). A conversation is a durable record of the message history stored in Foundry. With the Responses protocol the platform manages conversation history for you; with Invocations you manage session state in your own code. Session lifecycle. Compute is created on first use. After 15 minutes of inactivity the platform deprovisions the compute and persists the session state, then restores it automatically when the session resumes. A session is permanently deleted after 30 days of inactivity. This is where scale-to-zero comes from. Two identities. Each agent gets its own Microsoft Entra ID (agent identity) for runtime calls (models, tools, downstream services). Separately, the project managed identity is used by the platform for infrastructure operations like pulling the container image. You do not wire these up manually. Protocols. Hosted agents can speak Responses (OpenAI-compatible, platform-managed history and streaming), Invocations (arbitrary JSON in and out, you own the schema), and Invocations (WebSocket) for real-time voice. A single agent can expose more than one. Not sure which to pick? Start with Responses, you can add Invocations later. Sandbox sizes. You choose the CPU and memory per session: 0.5 vCPU / 1 GiB, 1 vCPU / 2 GiB, or 2 vCPU / 4 GiB. Billing is based on CPU and memory consumed across active sessions, so oversizing multiplies cost by your concurrency. Scaling. Hosted agents scale per session, not per replica: the platform spins up one isolated sandbox per session on demand and tears it down when the session ends. There\u0026rsquo;s no replica count or warm pool to size, and no fixed ceiling you configure, so capacity just follows your concurrent-session count and drops to zero when everything is idle. Requirements # Here is everything you need to follow along. The flow leans on the Azure Developer CLI (azd), which scaffolds the project and provisions the Azure side for you, so the list is shorter than it used to be.\nLocal tooling # An Azure subscription with access to Microsoft Foundry. Python 3.13+ (the azd ai agent run tooling and the supported hosting runtimes require 3.13 or newer). The Azure CLI and the Azure Developer CLI (azd) 1.25.3+. The Foundry extension for azd, which adds the azd ai agent ... commands: azd ext install microsoft.foundry. Confirm it with azd ext list. NOTE: Docker Desktop is not required. The image is always built remotely, so you don\u0026rsquo;t need a local Docker install for either deploy mode. The only difference is whether you author a Dockerfile: the default code mode builds the image from your source and declared runtime for you, while container mode uses your own Dockerfile for full control over the build.\nAzure resources # Under the hood a hosted agent needs a Foundry resource and project in a region that supports hosted agents and a model deployment (the wizard defaults to gpt-5.4-mini). In the default code + remote build mode that is essentially all that lands in your resource group. There is no container registry (the platform builds your image from the uploaded source and stores it on Microsoft-managed infrastructure); an Azure Container Registry only enters the picture on the container path, where you supply your own Dockerfile. Neither scaffold provisions Application Insights, so tracing is a quick opt-in: hit Connect in the portal\u0026rsquo;s Traces tab (or add it to the infra and re-provision) and traces start flowing. azd provision stands the rest up for you, no resources to create by hand.\nTIP: Prefer to own the infrastructure yourself (fixed names, an existing project, Bicep in source control)? You can provision separately and point azd at the existing project. My Foundry advanced workshop does exactly that with a main.bicep if you want that level of control.\nPermissions # Because azd provision creates the platform role assignments for you, there is just one you add by hand (your own). It still helps to picture the principals involved, three identities, each needing a role at a specific scope:\nflowchart LR You[Youuser] --\u003e|Owner at RG for a new projector Foundry Project Manager for an existing one| Proj[Foundry project] You -.-\u003e|Foundry Useradd by hand, needed for azd ai agent run and azd deploy| Acct[Foundry resource] ProjMI[Project managed identity] --\u003e|Container Registry Repository Readercontainer mode only| ACR[(Container registry)] ProjMI --\u003e|Foundry User| Acct Agent[Agent identitycreated at first deploy] --\u003e|Foundry User| Acct What you actually need on your own account depends on whether azd creates a new project or reuses an existing one, straight from the official permissions reference:\nCreating a new Foundry project: Owner at the resource group scope, so provision can create the resources and assign the roles below. Using an existing project: Foundry Project Manager at the project scope. Most of it is automatic. The project managed identity gets Foundry User on the Foundry resource (to run model inference), plus Container Registry Repository Reader on a registry when one is involved (the container path only), and each agent identity gets Foundry User when the platform creates it on the first deploy. The one assignment you do have to add by hand is Foundry User for your own identity: your CLI needs it for both azd ai agent run and azd deploy, since each acts as you against the project\u0026rsquo;s data plane. I cover it in Lesson 01.\nSign in # Authenticate once for both CLIs, az for your local credentials and azd for provisioning and deployment:\naz login azd auth login TIP: run azd config set auth.useAzCliAuth true once and azd reuses your az credential, so a single az login covers both.\nYou don\u0026rsquo;t need to hand-build a virtual environment or a .env for the azd flow; azd ai agent run handles the venv, dependencies, and environment values for you. (Running python main.py directly still reads a local .env.)\nCoding # Let\u0026rsquo;s build from the ground up. I will use Microsoft Agent Framework for the first two examples, then show how the exact same platform hosts a completely different framework in the third.\nLesson 01: your first hosted agent # The smallest possible hosted agent is genuinely small. You create a FoundryChatClient, wrap an Agent in a ResponsesHostServer, and call run():\nimport os from azure.identity import DefaultAzureCredential from dotenv import load_dotenv from agent_framework import Agent from agent_framework.foundry import FoundryChatClient from agent_framework_foundry_hosting import ResponsesHostServer load_dotenv() credential = DefaultAzureCredential() # --- Foundry client for the agent to use --- client = FoundryChatClient( project_endpoint=os.environ[\u0026#34;FOUNDRY_PROJECT_ENDPOINT\u0026#34;], model=os.environ[\u0026#34;AZURE_AI_MODEL_DEPLOYMENT_NAME\u0026#34;], credential=credential, ) # --- Agent definition and instructions --- agent = Agent( client=client, instructions=( \u0026#34;You are a helpful healthcare assistant. \u0026#34; \u0026#34;You answer questions about general health, wellness, and medical terminology. \u0026#34; \u0026#34;Always remind the user that your answers are for informational purposes only \u0026#34; \u0026#34;and not a substitute for professional medical advice.\u0026#34; ), default_options={\u0026#34;store\u0026#34;: False}, ) # --- Start the server to host the agent --- server = ResponsesHostServer(agent) server.run() The nice detail here is DefaultAzureCredential. Locally it uses your az login session. Once deployed, Foundry injects a managed identity automatically, so the exact same code authenticates in the cloud with no changes.\nAlongside main.py the only other file you bring is a requirements.txt with your dependencies. For this Agent Framework agent that is just the Foundry client and hosting packages (plus debugpy, so the Foundry Toolkit can attach a local debugger):\nagent-framework-foundry agent-framework-foundry-hosting\u0026gt;=1.0.0a260630 # debugpy enables local debugging of this agent with the Foundry Toolkit VS Code extension. debugpy azure-identity and python-dotenv, which main.py imports, come in transitively with agent-framework-foundry, so you don\u0026rsquo;t have to list them yourself.\nThere\u0026rsquo;s no hosting manifest to hand-write. When you run azd ai agent init (next step), the wizard generates an azure.yaml that declares your agent as a service, with the protocol and the model deployment name (AZURE_AI_MODEL_DEPLOYMENT_NAME) in that azure.ai.agent entry. The project endpoint needs no mapping: azd provision publishes it as FOUNDRY_PROJECT_ENDPOINT and the platform injects it at runtime.\nNOTE: those variable names are a convention, not something the tooling scrapes from your code. The sample\u0026rsquo;s .env.example spells out the contract; azd provision fills both in (writing them into .azure/\u0026lt;env\u0026gt;/.env) and injects them at runtime, so the names your main.py reads must match:\nFOUNDRY_PROJECT_ENDPOINT=\u0026#34;...\u0026#34; AZURE_AI_MODEL_DEPLOYMENT_NAME=\u0026#34;...\u0026#34; That is essentially the whole agent, no Dockerfile needed in the default code mode. (Prefer to own the build? Use --deploy-mode container with your own Dockerfile on port 8088.)\nWith the tooling from the Requirements section in place, the loop is short. Put your main.py and requirements.txt in a folder, then walk it one step at a time.\n1. Initialize the project:\nazd ai agent init --deploy-mode code Because the folder already contains your code, azd ai agent init offers Use the code in the current directory. Say yes, then answer the wizard: a project and agent name, the runtime (Python 3.13), the entry point (main.py), how dependencies are resolved (Remote build), the protocol (responses), whether to Create a new Foundry project or reuse an existing one, then your subscription, region, and model (deploy a new one from the catalog, the wizard suggests gpt-5.4-mini). When it finishes you have an azure.yaml with your agent wired up and an environment seeded under .azure/\u0026lt;env\u0026gt;/, no files to copy, no manifest to hand-write.\nTIP: prefer a running starting point over your own blank main.py? Initialize in an empty folder instead and the wizard downloads the ready-made Basic Agent Framework sample (into a src/ subfolder), then replace its main.py with your own.\n2. Provision the Azure resources:\nazd provision This is the step that does the heavy lifting: it reads azure.yaml, stands up the Azure resources (on the code path that is the Foundry project and model deployment), and writes their values (endpoint, model, project ID) straight into the azd environment. No manual variable juggling, and the platform\u0026rsquo;s role assignments are created for you. That leaves just one role for you to add by hand, and the flow is still far shorter than wiring everything up yourself.\nNOTE: that role is for your own identity. Your CLI performs data-plane actions against the project: azd ai agent run executes main.py locally and calls the model through the project endpoint, and azd deploy creates the agent version (an agents/write action). Both run as you, not a managed identity. Owner and Contributor are control-plane roles and don\u0026rsquo;t cover these, so you can hit a 403, either \u0026ldquo;You don\u0026rsquo;t have permission to build agents in this project\u0026rdquo; on a local run or \u0026ldquo;does not have permissions for \u0026hellip;agents/write\u0026rdquo; on deploy. The fix is the least-privilege data-plane role, Foundry User. The Foundry portal even offers an Assign me the Foundry User role button. From the CLI, assign it at the Foundry resource scope after azd provision:\n# Your signed-in user\u0026#39;s object ID az ad signed-in-user show --query id -o tsv # Grant the Foundry User data-plane role at the Foundry resource (account) scope. az role assignment create \\ --assignee \u0026lt;your-object-id\u0026gt; \\ --role \u0026#34;Foundry User\u0026#34; \\ --scope /subscriptions/\u0026lt;subscription-id\u0026gt;/resourceGroups/\u0026lt;resource-group\u0026gt;/providers/Microsoft.CognitiveServices/accounts/\u0026lt;foundry-account\u0026gt; Give it a minute or two to propagate, then re-run the command. This applies to both the code and container modes (I hit the deploy 403 on the container path). What the managed identity covers is the deployed agent at runtime, calling models from inside its own project, not your CLI commands. For the full role matrix, see the hosted agent permissions reference.\n3. Run the agent locally:\nazd ai agent run This builds a virtual environment, installs the dependencies, and starts the agent, but it doesn\u0026rsquo;t stop there. It also pops open a browser with the Agent Inspector, a chat UI wired straight to your local http://localhost:8088 endpoint. You can send prompts and watch the raw Responses API events stream in on the side (response.output_text.delta, response.completed, and friends), which makes it easy to see exactly what the agent is emitting. Great for a quick sanity check before you deploy.\nTIP: azd ai agent invoke --local \u0026quot;...\u0026quot; hits that same local endpoint from a second terminal, so you can script quick checks while the Inspector is open.\n4. Deploy to Foundry:\nazd deploy The platform builds the container image remotely from your source and declared runtime and ships it to Foundry as a hosted agent, storing it on managed infrastructure (on the code path there\u0026rsquo;s no registry in your resource group). No local Docker, no image to babysit.\n5. Invoke the deployed agent:\nazd ai agent invoke \u0026#34;What are the symptoms of vitamin D deficiency?\u0026#34; This calls the deployed endpoint (now running under the project managed identity) and streams the response back to your terminal. That is the full loop, from your own main.py to a running hosted agent in Foundry.\nLesson 02: tools and file persistence # An agent that only chats is not that interesting. The feature I really wanted to try here is per-session file persistence: each session gets a VM-isolated sandbox with a persistent filesystem, so anything the agent writes to disk survives idle periods and is restored when the session resumes. No database, no blob storage, no wiring. To reach that filesystem the agent needs tools, and with Agent Framework a tool is just a plain Python function you decorate with @tool.\nLet\u0026rsquo;s give the agent two tools that read and write notes on the sandbox. The docs list $HOME (and /files) as the persistent locations, so I write under $HOME/notes. Drop these two functions into your main.py from Lesson 01, above the agent definition:\n# --- New imports --- from datetime import datetime from typing import Annotated from agent_framework import tool from pydantic import Field NOTES_DIR = os.path.join(os.path.expanduser(\u0026#34;~\u0026#34;), \u0026#34;notes\u0026#34;) # persistent per-session $HOME # --- Tools for the agent to use --- @tool(approval_mode=\u0026#34;never_require\u0026#34;) def save_session_note( note: Annotated[str, Field(description=\u0026#34;The note text to save\u0026#34;)], ) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Save a note to the per-session sandbox filesystem.\u0026#34;\u0026#34;\u0026#34; os.makedirs(NOTES_DIR, exist_ok=True) timestamp = datetime.now().strftime(\u0026#34;%Y%m%d_%H%M%S\u0026#34;) filepath = os.path.join(NOTES_DIR, f\u0026#34;note_{timestamp}.txt\u0026#34;) with open(filepath, \u0026#34;w\u0026#34;) as f: f.write(note) return f\u0026#34;Note saved to {filepath}\u0026#34; @tool(approval_mode=\u0026#34;never_require\u0026#34;) def list_session_notes() -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;List all notes saved in the current session.\u0026#34;\u0026#34;\u0026#34; if not os.path.exists(NOTES_DIR): return \u0026#34;No notes found.\u0026#34; files = sorted(os.listdir(NOTES_DIR)) if not files: return \u0026#34;No notes found.\u0026#34; results = [] for fname in files: with open(os.path.join(NOTES_DIR, fname)) as f: results.append(f\u0026#34;--- {fname} ---\\n{f.read()}\u0026#34;) return \u0026#34;\\n\\n\u0026#34;.join(results) Then hand both tools to the agent. This replaces the Agent(...) block from Lesson 01, everything else (the FoundryChatClient and the ResponsesHostServer(agent).run() at the bottom) stays exactly the same:\n# --- Agent definition with tools and instructions --- agent = Agent( client=client, instructions=( \u0026#34;You are a helpful healthcare assistant with access to health tools. \u0026#34; \u0026#34;Use save_session_note and list_session_notes to manage the user\u0026#39;s notes. \u0026#34; \u0026#34;Always remind the user that your answers are for informational purposes only \u0026#34; \u0026#34;and not a substitute for professional medical advice.\u0026#34; ), tools=[save_session_note, list_session_notes], default_options={\u0026#34;store\u0026#34;: False}, ) NOTE: the $HOME sandbox only exists on the hosted platform. Running these tools under azd ai agent run writes to your laptop\u0026rsquo;s home directory, which tells you nothing about session persistence, so test this one after azd deploy.\nDeploy, then save a note and list it back in two separate calls:\nazd ai agent invoke \u0026#34;Save a note: Patient P-1001 follow-up scheduled for December.\u0026#34; azd ai agent invoke \u0026#34;List my session notes\u0026#34; The note is still there on the second call, even though the compute scaled down in between and the session resumed a fresh sandbox. You did not provision a database or wire up storage; the platform did the boring-but-critical work for you.\nTIP: consecutive invokes reuse the last session automatically, so the two calls above land in the same sandbox. Pass --new-session to start fresh, or pin a specific session with --session-id \u0026lt;id\u0026gt; (the ID is printed on every invoke). Make sure you are on a recent azure.ai.agents extension; on an older beta the auto-reuse misfired for me and every invoke got a new empty session.\nLesson 03: bring your own framework (LangGraph) # Here is the part I find genuinely powerful. Foundry does not care which framework you use. It only needs your agent to speak the Responses protocol on port 8088. So you can take a completely different stack, in this case LangGraph, and host it on the exact same platform: same azd flow, same per-session sandbox, same managed identity.\nThis used to require a hand-written adapter to translate Responses requests into LangChain messages and back. That is no longer necessary. Microsoft now ships a supported hosting package, langchain-azure-ai[hosting], whose ResponsesHostServer takes a compiled LangGraph graph and handles all the protocol plumbing for you. You build the graph, it does the rest.\nHere is the whole agent: the same healthcare assistant, this time with two new tools (patient lookup and BMI), built with LangChain\u0026rsquo;s @tool, a LangGraph StateGraph, and the official host. This is a fresh project, so it is a complete main.py, not a diff:\n# --- Lesson 03 — LangGraph hosted agent (bring your own framework) --- import json import os from typing import Annotated, Any from azure.identity import DefaultAzureCredential from dotenv import load_dotenv from langchain_azure_ai.agents.hosting import ResponsesHostServer from langchain_azure_ai.chat_models import AzureAIOpenAIApiChatModel from langchain_core.messages import AIMessage, SystemMessage from langchain_core.tools import tool from langgraph.graph import END, START, StateGraph from langgraph.graph.message import add_messages from langgraph.prebuilt import ToolNode from pydantic import Field from typing_extensions import TypedDict load_dotenv() credential = DefaultAzureCredential() # --- Tools --- @tool def lookup_patient_record( patient_id: Annotated[str, Field(description=\u0026#34;Patient ID, e.g. P-1001\u0026#34;)], ) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Look up a patient record by ID.\u0026#34;\u0026#34;\u0026#34; records = { \u0026#34;P-1001\u0026#34;: {\u0026#34;name\u0026#34;: \u0026#34;Alice Johnson\u0026#34;, \u0026#34;age\u0026#34;: 34, \u0026#34;blood_type\u0026#34;: \u0026#34;A+\u0026#34;, \u0026#34;conditions\u0026#34;: [\u0026#34;asthma\u0026#34;]}, \u0026#34;P-1002\u0026#34;: {\u0026#34;name\u0026#34;: \u0026#34;Bob Martinez\u0026#34;, \u0026#34;age\u0026#34;: 58, \u0026#34;blood_type\u0026#34;: \u0026#34;O-\u0026#34;, \u0026#34;conditions\u0026#34;: [\u0026#34;type 2 diabetes\u0026#34;, \u0026#34;hypertension\u0026#34;]}, } record = records.get(patient_id) if record is None: return f\u0026#34;No patient found with ID {patient_id}.\u0026#34; return json.dumps(record, indent=2) @tool def calculate_bmi( weight_kg: Annotated[float, Field(description=\u0026#34;Weight in kilograms\u0026#34;)], height_m: Annotated[float, Field(description=\u0026#34;Height in meters\u0026#34;)], ) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Calculate Body Mass Index (BMI) from weight and height.\u0026#34;\u0026#34;\u0026#34; if height_m \u0026lt;= 0: return \u0026#34;Height must be greater than zero.\u0026#34; bmi = weight_kg / (height_m ** 2) category = ( \u0026#34;underweight\u0026#34; if bmi \u0026lt; 18.5 else \u0026#34;normal weight\u0026#34; if bmi \u0026lt; 25 else \u0026#34;overweight\u0026#34; if bmi \u0026lt; 30 else \u0026#34;obese\u0026#34; ) return f\u0026#34;BMI: {bmi:.1f} ({category})\u0026#34; tools = [lookup_patient_record, calculate_bmi] # --- Model (Foundry deployment via langchain-azure-ai) --- llm = AzureAIOpenAIApiChatModel( project_endpoint=os.environ[\u0026#34;FOUNDRY_PROJECT_ENDPOINT\u0026#34;], model=os.environ[\u0026#34;AZURE_AI_MODEL_DEPLOYMENT_NAME\u0026#34;], credential=credential, ).bind_tools(tools) # --- LangGraph graph --- SYSTEM_PROMPT = ( \u0026#34;You are a helpful healthcare assistant. \u0026#34; \u0026#34;You can look up patient records and calculate BMI. \u0026#34; \u0026#34;Always remind the user your answers are for informational purposes only.\u0026#34; ) class AgentState(TypedDict): messages: Annotated[list, add_messages] def chatbot(state: AgentState) -\u0026gt; dict[str, Any]: messages = [SystemMessage(content=SYSTEM_PROMPT)] + state[\u0026#34;messages\u0026#34;] return {\u0026#34;messages\u0026#34;: [llm.invoke(messages)]} def should_continue(state: AgentState) -\u0026gt; str: last_message = state[\u0026#34;messages\u0026#34;][-1] if isinstance(last_message, AIMessage) and last_message.tool_calls: return \u0026#34;tools\u0026#34; return END graph_builder = StateGraph(AgentState) graph_builder.add_node(\u0026#34;chatbot\u0026#34;, chatbot) graph_builder.add_node(\u0026#34;tools\u0026#34;, ToolNode(tools)) graph_builder.add_edge(START, \u0026#34;chatbot\u0026#34;) graph_builder.add_conditional_edges(\u0026#34;chatbot\u0026#34;, should_continue, {\u0026#34;tools\u0026#34;: \u0026#34;tools\u0026#34;, END: END}) graph_builder.add_edge(\u0026#34;tools\u0026#34;, \u0026#34;chatbot\u0026#34;) graph = graph_builder.compile() # --- Serve on the Foundry Responses protocol --- if __name__ == \u0026#34;__main__\u0026#34;: port = int(os.environ.get(\u0026#34;PORT\u0026#34;, \u0026#34;8088\u0026#34;)) ResponsesHostServer(graph).run(port=port) A few things worth pointing out:\nThe tools use LangChain\u0026rsquo;s @tool decorator (note there is no approval_mode here, that is an Agent Framework detail). AzureAIOpenAIApiChatModel connects to your Foundry model deployment, and .bind_tools(tools) tells the model what it can call. It reads the same FOUNDRY_PROJECT_ENDPOINT and AZURE_AI_MODEL_DEPLOYMENT_NAME values as Lesson 01. The StateGraph is the part you own: a chatbot node calls the model, a tools node runs any tool calls, and should_continue loops back until the model is done. That explicit loop is exactly what Agent Framework handles internally for you, here you get to see and shape it. ResponsesHostServer(graph).run(port=port) is the entire hosting layer. Because a normal LangGraph graph already keeps its state in a messages field, the host wires straight into it, no translation code. The requirements.txt pulls the hosting package and LangGraph:\nlangchain-azure-ai[hosting]\u0026gt;=1.2.4 langgraph langchain-core azure-identity python-dotenv From here the flow is identical to Lesson 01. Point azd ai agent init at this folder, reuse your existing Foundry project and model deployment from the previous lessons, then provision and deploy:\nazd ai agent init --deploy-mode code azd provision azd deploy Then invoke it with a prompt that needs both tools in a single turn:\nazd ai agent invoke \u0026#34;Look up patient P-1002 and calculate BMI for 95kg at 1.80m\u0026#34; The agent comes back with Bob Martinez\u0026rsquo;s record and a BMI of 29.3 (overweight), the LangGraph loop deciding on its own to call both tools before answering. Same platform, same commands, a completely different framework under the hood.\nThe azure.yaml service entry is identical to the previous lessons (and if you use container mode, so is the Dockerfile). That is the whole point: no lock-in. Foundry is multi-model and multi-harness by design, so you can run models from OpenAI, Anthropic, Meta, Mistral and others, and bring LangGraph, Agent Framework, the OpenAI Agents SDK, or your own custom loop.\nWhat\u0026rsquo;s new since the preview # Because this space moves fast, it is worth calling out the updates Microsoft shipped after the initial refresh. A few stood out to me:\nDeploy straight from source code, no container required. This is the code deploy mode we used above: you point azd at a Python or .NET project and let the platform build the image, with no Dockerfile to maintain. You can also scaffold it directly from existing source:\nazd ai agent init \\ --src ./src/my-agent \\ --agent-name my-unique-agent \\ --deploy-mode code \\ --runtime python_3_13 \\ --entry-point main.py \\ --dep-resolution remote_build azd deploy Supported runtimes are python_3_13, python_3_14, and dotnet_10. Container-based deployment is still fully supported when you want full control over the runtime. The hosted agent quickstart (azd) walks through the source-code flow end to end.\nBuilt-in content safety guardrails. Instead of standing up your own Content Safety endpoint and middleware, you can attach guardrail policies to the agent definition. Every prompt is checked before it reaches your code, and every response before it reaches the user.\nVoice Live and WebSocket support. The Invocations (WebSocket) protocol adds a persistent bidirectional connection for real-time voice agents, pairing with frameworks like Voice Live, Pipecat, or LiveKit. It started out limited to North Central US, and per the docs the WebSocket protocol is now available across all regions that support hosted agents.\nAgent optimizer. A closed-loop engine that evaluates your agent against defined criteria, generates better instructions (or skills, model choices, and tool descriptions), ranks the candidates, and lets you deploy the winner as a new version. It is now in public preview.\nClosing # What I like most about hosted agents is that they let me keep the fun part (writing the agent) and hand off the hard part (running it safely at scale). Per-session isolation, a persistent filesystem, a built-in identity, scale-to-zero economics, and observability out of the box are exactly the things I used to cobble together by hand. Add source-code deployment, guardrails, and the agent optimizer on top, and the path from \u0026ldquo;works on my machine\u0026rdquo; to \u0026ldquo;running in production\u0026rdquo; gets dramatically shorter.\nIf you want to get your hands dirty, the Microsoft Foundry Advanced Workshop walks through all of this end to end, from your first hosted agent to Toolbox, guardrails, and knowledge grounding. And since this space keeps evolving fast, keep an eye on the official docs for the latest. Happy building!\nSources # What are hosted agents? (Microsoft Learn) Introducing the new hosted agents in Foundry Agent Service (Microsoft Foundry Blog) What\u0026rsquo;s New in Hosted Agents in Foundry Agent Service (Microsoft Foundry Blog) Microsoft Foundry Advanced Workshop Microsoft Agent Framework documentation Host LangGraph agents as Foundry hosted agents What is Microsoft Foundry? Hosted agents region availability Install the Azure CLI Install the Azure Developer CLI (azd) Download Python Create an Azure free account ","date":"22 July 2026","externalUrl":null,"permalink":"/posts/hosted-agents/","section":"Posts","summary":"Intro # Building an agent on your laptop is the easy part.","title":"Hosted Agents in Microsoft Foundry: Bring Your Own Code, Ship to Production","type":"posts"},{"content":"","date":"22 July 2026","externalUrl":null,"permalink":"/tags/langgraph/","section":"Tags","summary":"","title":"LangGraph","type":"tags"},{"content":"","date":"17 July 2026","externalUrl":null,"permalink":"/tags/mcp/","section":"Tags","summary":"","title":"MCP","type":"tags"},{"content":" Intro # The platform formerly known as Azure AI Foundry has quietly turned into something quite different over the course of this year: a new name, a new portal, and a genuinely new engine under the hood. Welcome to Microsoft Foundry and the V2 wave of changes. None of this landed overnight, it has been rolling out in stages since the beginning of the year, for example the Foundry Agent Service went GA back in March. By now enough of it has settled that it is worth stepping back and looking at the whole picture. If you have read my earlier post Azure AI Foundry News \u0026amp; Changes, this is the sequel: what was in preview back then has matured, the rebrand is official, and the way we build agents has changed quite a bit.\nIn this post I want to give you a quick overview of the most important changes first, then dive into the details with code examples: the Responses API, conversations replacing threads, prompt agents, the growing set of Foundry tools, the new portal UI, and how you operate and monitor everything. I will also call out what changed compared to my previous Foundry post, because the shift from threads and runs to conversations and responses is the part that will bite you if you copy old code. Let\u0026rsquo;s check it out.\nNOTE: A lot of this space is moving fast, and some capabilities are still in preview. I link to the official docs throughout so you always have the current state of GA vs preview.\nOverview: the big changes # Let\u0026rsquo;s start with the headline changes before we get into code. The single most important thing to internalize is that this is not just a rename, it is a new API generation. The official What is Microsoft Foundry? doc has a great mapping table from the old world to the new one. Here is the short version:\nDimension Previous (classic) Current (V2) Brand Azure AI Studio / Azure AI Foundry Microsoft Foundry Azure AI Services Azure AI Services Foundry Tools Portal Foundry (classic) Foundry (new portal) Agent API Assistants API (Agents v0.5/v1) Responses API (Agents v2) API versioning Monthly api-version params v1 stable routes (/openai/v1/) Resource model Hub + Azure OpenAI + Azure AI Services Foundry resource (single, with projects) SDKs \u0026amp; endpoints Multiple packages against 5+ endpoints Unified project client (azure-ai-projects 2.x) + OpenAI() Terminology Threads, Messages, Runs, Assistants Conversations, Items, Responses, Agent Versions If you only remember one row, make it the last one. Threads, messages, runs, and assistants are out. Conversations, items, responses, and agent versions are in. Everything now builds on the Responses API, the same modern primitive you may know from OpenAI, exposed through stable /openai/v1/ routes so you are no longer chasing monthly api-version strings.\nThe other big story is the resource simplification I teased in my last post. In the classic Hub-based world you needed a whole pile of Azure resources. The new model is a single Foundry resource with projects, and the old Azure AI Services brand is now called Foundry Tools. One resource provider namespace, unified RBAC, networking, and policies.\nTIP: Coming from a plain Azure OpenAI resource? You can upgrade it to a Foundry resource and keep your endpoint, keys, and existing state.\nWhat changed since my last Foundry post # In my previous Foundry post I was still using threads, messages, and runs, and there was no Responses API in sight. Here is the exact code I wrote back then to talk to an agent:\n# The old way: threads + messages + runs thread = project_client.agents.threads.create() message = project_client.agents.messages.create( thread_id=thread.id, role=\u0026#34;user\u0026#34;, content=\u0026#34;Who is the greatest Basketball player of all time?\u0026#34;, ) run = project_client.agents.runs.create_and_process( thread_id=thread.id, agent_id=agent.id, ) In V2 that same interaction becomes a conversation plus a response. Runs with their queued/in-progress polling loops are gone, and you talk to the model through the OpenAI-compatible client that the project hands you:\n# The new way: conversations + responses conversation = openai.conversations.create( items=[ { \u0026#34;type\u0026#34;: \u0026#34;message\u0026#34;, \u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;Who is the greatest Basketball player of all time?\u0026#34;, } ], ) response = openai.responses.create( conversation=conversation.id, input=\u0026#34;Who is the greatest Basketball player of all time?\u0026#34;, extra_body={ \u0026#34;agent_reference\u0026#34;: {\u0026#34;name\u0026#34;: agent.name, \u0026#34;type\u0026#34;: \u0026#34;agent_reference\u0026#34;}, }, ) Microsoft published a detailed Migrate to the new Foundry Agent Service guide that walks through threads to conversations, runs to responses, and assistants to new agents. There is even a migration tool that automates the code constructs for you. One caveat worth repeating from the docs: the migration tool moves your code, not your state. Old threads, runs, and messages are not migrated, so you start fresh conversations after moving over.\nHere is the mental model as a diagram:\nflowchart LR subgraph Classic[\"Classic (v1)\"] A1[Assistant] --\u003e A2[Thread] A2 --\u003e A3[Messages] A3 --\u003e A4[Run + polling] end subgraph V2[\"Microsoft Foundry V2\"] B1[Agent version] --\u003e B2[Conversation] B2 --\u003e B3[Items] B3 --\u003e B4[Response] end Classic --\u003e|migrate| V2 Setup and the Responses API # Let\u0026rsquo;s build this up from scratch. First, grab the SDK. The V2 surface lives in azure-ai-projects 2.x (it pulls in the openai package as a dependency), and you\u0026rsquo;ll want azure-identity for DefaultAzureCredential:\npip install \u0026#34;azure-ai-projects\u0026gt;=2.3.0\u0026#34; azure-identity As before, you connect with a project endpoint (no more connection strings) and DefaultAzureCredential. The new twist is that the project client hands you an OpenAI-compatible client via get_openai_client(), and that is what you use for conversations and responses:\nfrom azure.identity import DefaultAzureCredential from azure.ai.projects import AIProjectClient # Format: https://\u0026lt;resource\u0026gt;.services.ai.azure.com/api/projects/\u0026lt;project\u0026gt; PROJECT_ENDPOINT = \u0026#34;your_project_endpoint\u0026#34; project = AIProjectClient( endpoint=PROJECT_ENDPOINT, credential=DefaultAzureCredential(), ) # The OpenAI-compatible client for conversations and responses openai = project.get_openai_client() The simplest thing you can do is a plain model call through the Responses API, no agent required:\nresponse = openai.responses.create( model=\u0026#34;gpt-5-mini\u0026#34;, input=\u0026#34;What are the common symptoms of type 2 diabetes?\u0026#34;, ) print(response.output_text) That is the whole point of V2: one project endpoint, one OpenAI-compatible client for every model, one Responses API. The same responses.create() call is used whether you hit a raw model or route through an agent. You can even let Foundry pick the model for you with model routing.\nNOTE: The division of labor trips people up at first. Use the project client for agent creation and versioning, and the OpenAI client (project.get_openai_client()) for conversations and responses. If you call conversations.create() on the project client, it will not exist.\nConversations instead of threads # A conversation is the successor to a thread, but it stores a stream of items, not just messages. Items can be messages, tool calls, tool outputs, and more. You create one, then keep sending responses against it, and it retains context across calls automatically. Here is a complete two-turn exchange, no agent required, just a model and the conversation:\n# The conversation holds the running state conversation = openai.conversations.create( metadata={\u0026#34;agent\u0026#34;: \u0026#34;patient-education-agent\u0026#34;}, ) # First turn: the response is appended to the conversation automatically response = openai.responses.create( model=\u0026#34;gpt-5-mini\u0026#34;, conversation=conversation.id, input=\u0026#34;Explain what an mRNA vaccine is in one sentence.\u0026#34;, ) print(response.output_text) # Follow-up turn: notice we never repeat the earlier context, # the conversation remembers it for us response = openai.responses.create( model=\u0026#34;gpt-5-mini\u0026#34;, conversation=conversation.id, input=\u0026#34;How is it different from a traditional vaccine?\u0026#34;, ) print(response.output_text) That second question only makes sense because the conversation carried the context from the first turn. A multi-turn chat is just repeated responses.create() calls that reference the same conversation.id, no more manual message list juggling. If you ever need to seed a conversation with prior items without generating a response, you can add them directly with openai.conversations.items.create(conversation_id=..., items=[...]).\nPrompt agents # Agents themselves got a redesign too. Instead of create_agent(), you now create versions of an agent with a structured definition. The most common type is a prompt agent: a declarative agent defined by a model plus instructions (and optionally tools). Here is the pattern straight from the migration guide and my Foundry workshop:\nfrom azure.ai.projects.models import CodeInterpreterTool, PromptAgentDefinition agent = project.agents.create_version( agent_name=\u0026#34;lab-results-agent\u0026#34;, definition=PromptAgentDefinition( model=\u0026#34;gpt-5-mini\u0026#34;, instructions=( \u0026#34;You politely help interpret patient lab results. Use the Code \u0026#34; \u0026#34;Interpreter tool when asked to visualize biomarker trends.\u0026#34; ), tools=[CodeInterpreterTool()], ), ) Note that the instructions reference the Code Interpreter tool, but the guidance alone is not enough, you also have to attach the tool via tools=[CodeInterpreterTool()]. It is a built-in tool, so Foundry runs it server-side in a sandbox and you do not need to host anything or add a connection. With the tool attached, the plot request below can actually execute Python and return a chart.\nNotice the two nice properties this gives you. First, versioning is built in: every call to create_version() under the same agent_name produces a new version (lab-results-agent:1, lab-results-agent:2, and so on), so you can iterate without deleting the old one. Second, there is a clean separation of duties: you define the agent once, then execute it with different inputs by passing an agent_reference on your responses.\nTo actually talk to the agent, you combine everything from above: a conversation for state, and a response that references the agent:\n# A fresh conversation to hold the agent\u0026#39;s working state conversation = openai.conversations.create() response = openai.responses.create( conversation=conversation.id, input=( \u0026#34;Plot this patient\u0026#39;s HbA1c readings over the last six months: \u0026#34; \u0026#34;7.8, 7.4, 7.1, 6.9, 6.7, 6.5.\u0026#34; ), extra_body={ \u0026#34;agent_reference\u0026#34;: {\u0026#34;name\u0026#34;: agent.name, \u0026#34;type\u0026#34;: \u0026#34;agent_reference\u0026#34;}, }, ) for item in response.output: if item.type == \u0026#34;message\u0026#34;: for block in item.content: print(block.text) The loop above prints the agent\u0026rsquo;s text explanation. The chart itself comes back as its own output item (Code Interpreter returns files/images separately, not inside output_text), so in a real app you would also iterate the non-message items to grab the generated image.\nThere is a second agent type, the Hosted agent, for running your own agent code and containers on Foundry. That deserves its own hands-on walkthrough, so I will cover Hosted agents in a separate post. For this one, prompt agents are all we need.\nFoundry tools # Tools are where V2 really opens up. The new Foundry Agent Service ships a broad set of built-in tools, some carried over from classic and some brand new. A few highlights from the migration guide\u0026rsquo;s tool availability table:\nTool Classic V2 Web Search No Yes (GA) Image Generation No Yes (Public Preview) MCP Public Preview Yes (GA) Agent to Agent (A2A) No Yes (Public Preview) Code Interpreter Yes (GA) Yes (GA) File Search Yes (GA) Yes (GA) Azure AI Search Yes (GA) Yes (GA) Grounding with Bing Search Yes (GA) Yes (GA) OpenAPI Yes (GA) Yes (GA) Two things jump out for me: MCP is now GA (it was preview in my last post), and Web Search and Image Generation are new. Beyond the built-ins, Foundry now has a tool catalog with over 1,400 tools via public and private catalogs.\nAdding a tool to a prompt agent is just a line in the definition. Here is the web search example from my workshop\u0026rsquo;s tool calls lab, and the official web search tool docs have the full reference:\nfrom azure.ai.projects.models import PromptAgentDefinition, WebSearchTool agent = project.agents.create_version( agent_name=\u0026#34;clinical-web-search-agent\u0026#34;, definition=PromptAgentDefinition( model=\u0026#34;gpt-5-mini\u0026#34;, instructions=\u0026#34;Answer medical questions and use web search for recent clinical guidelines.\u0026#34;, tools=[WebSearchTool()], ), ) When you send a time-sensitive question through responses.create(), the agent decides on its own to call the tool. The final answer still comes back as plain text, but the response payload also includes structured items describing what happened, so you can confirm the tool actually ran:\nresponse = openai.responses.create( input=\u0026#34;What are the latest FDA-approved treatments for rheumatoid arthritis?\u0026#34;, extra_body={ \u0026#34;agent_reference\u0026#34;: {\u0026#34;name\u0026#34;: agent.name, \u0026#34;type\u0026#34;: \u0026#34;agent_reference\u0026#34;}, }, ) # The answer itself is plain text print(response.output_text) # ...but the structured items tell you whether web search was actually used tool_used = any( getattr(item, \u0026#34;type\u0026#34;, \u0026#34;\u0026#34;) == \u0026#34;web_search_call\u0026#34; for item in response.output ) print(f\u0026#34;Web search tool used: {tool_used}\u0026#34;) TIP: A tool-backed answer and a plain model answer look identical to the user. The structured items in the response are how you verify grounding actually happened, which is gold for debugging and evaluation.\nThe new portal # The portal moved too. Everything now lives at ai.azure.com, and to get the V2 experience you make sure the New Foundry toggle in the banner is switched on. Foundry (classic) still exists for Hub-based projects, but all new investment is going into the new portal.\nIf you are used to the classic layout, some things moved around. Microsoft has a handy Find features in the Foundry portal guide to help you relocate everything. The playgrounds are still there for trying models and agents in the browser before you drop into code, and of course my favorite workflow, developing in VS Code, is fully supported through the Microsoft Foundry for VS Code extension.\nHere are the parts of the new UI I think are worth showing.\nThe new portal home at ai.azure.com with the \u0026ldquo;New Foundry\u0026rdquo; toggle switched on in the banner:\nThe model catalog under the \u0026ldquo;Discover\u0026rdquo; tab, to choose from over 11,000 different models:\nFoundry Tools, also under the \u0026ldquo;Discover\u0026rdquo; tab, where you browse the tool catalog and configure tools for your agents:\nThe Agents view (list of agents and agent versions) in the new portal under the \u0026ldquo;Build\u0026rdquo; tab:\nThe deployments view to list all your model deployments:\nThe playground for trying a model or agent in the browser:\nThe monitor agents dashboard with metrics, agent runs, and more:\nThe Operate overview dashboard under the \u0026ldquo;Operate\u0026rdquo; tab, with alerts, success rates, token usage, and cost:\nOperate and monitor # The part I appreciate most as a platform person is that operate and monitor is now first-class, not an afterthought. The new portal has a dedicated Operate section for centralized management of all your AI assets: agents, models, and tools in one place, including agents registered from other clouds.\nOn the observability side, Foundry gives you:\nNear-real-time observability with built-in operational metrics (token usage, latency, run success rate) surfaced in the monitor agents dashboard. It reads telemetry from a connected Application Insights resource, so expect a short ingestion delay rather than a live stream. Continuous evaluation so you keep scoring quality and safety on live traffic, not just in a one-off test run. Tracing built on OpenTelemetry, which you can wire up in a few lines and then inspect the traces right inside Foundry. I walk through exactly this in my workshop\u0026rsquo;s observability lab. Enterprise controls: full authentication for MCP and A2A, AI gateway integration, and Azure Policy support. That combination of build, operate, and govern under one resource is really what \u0026ldquo;V2\u0026rdquo; is about for me. It is less a fresh coat of paint and more a consolidation of the whole lifecycle.\nNOTE: a few of these pieces are still in public preview, such as the portal View agent metrics view, scheduled evaluations, red team scans, and alerts. Expect them to keep evolving.\nA few things I did not cover # Foundry V2 is broad, so I deliberately kept this post on the core building blocks. A few areas I skipped but that are well worth a look:\nPublishing agents where people already work. Once an agent is ready, you can push it straight to Microsoft 365 Copilot and Teams, so colleagues discover and chat with it without leaving their usual tools. What gets published is the agent\u0026rsquo;s stable endpoint, so you can roll out new versions behind the scenes. Multi-agent orchestration. For coordinating several specialized agents, the go-to engine is Microsoft Agent Framework, which I wrote a whole post on. The Foundry portal also has a visual Workflows designer, but it is being retired on December 1, 2026, so for anything new I would start with Agent Framework. Foundry IQ. A managed knowledge layer that connects agents to your enterprise data (Azure Blob Storage, SharePoint, OneLake, and the public web) through agentic retrieval. One knowledge base can serve many agents with permission-aware, citation-backed grounding, so you do not have to wire each agent to each source yourself. AI Services in the portal. It is not all LLMs. The portal also surfaces the classic Azure AI capabilities you can use straight from Foundry, like Speech (speech-to-text, text-to-speech, and even a photorealistic text-to-speech avatar), Vision, Language, and more. And that is still not everything, Foundry keeps growing, so treat this list as a starting point rather than the full map.\nWrap-up # So where does that leave us? Microsoft Foundry V2 is a meaningful step up from the Azure AI Foundry I wrote about last time. The rebrand is the least interesting part. The real changes are the Responses API as the single primitive, conversations and items replacing threads and messages, versioned prompt agents replacing assistants, a much richer tool story (MCP GA, Web Search, Image Generation), a new portal, and a proper operate and monitor experience.\nIf you have older Foundry code, do not just copy your threads-and-runs snippets, plan a small migration to conversations and responses. The migration guide and the migration tool make it painless. And if you want a hands-on, runnable path through all of this in Python, I put together a Microsoft Foundry workshop (GitHub repo) that covers model calls, prompt agents, tools, RAG, evaluation, observability, and more.\nNext up, I will dig into Hosted agents in a dedicated post, since running your own agent code on Foundry deserves more than a paragraph. Until then, flip on that New Foundry toggle and have fun building.\nSources # What is Microsoft Foundry? Migrate to the new Foundry Agent Service Foundry migration tool (GitHub) Microsoft Foundry SDKs and Endpoints Quickstart: Build with models and agents Model routing with the Responses API Agent tool catalog Use the web search tool in Foundry Agent Service Find features in the Foundry portal Monitor agents dashboard Microsoft Foundry for VS Code My previous post: Azure AI Foundry News \u0026amp; Changes Microsoft Foundry workshop and GitHub repo ","date":"17 July 2026","externalUrl":null,"permalink":"/posts/microsoft-foundry-v2/","section":"Posts","summary":"Intro # The platform formerly known as Azure AI Foundry has quietly turned into something quite different over the course of this year: a new name, a new portal, and a genuinely new engine under the hood.","title":"Microsoft Foundry V2: New Portal, Responses API, and Prompt Agents","type":"posts"},{"content":" Intro # The Microsoft Agent Framework has been around for a while now and, with the 1.0.0 release, it is officially generally available. What is it you say? You might have heard about the open-source frameworks Semantic Kernel and AutoGen for building agentic apps? Both have their strengths and their weaknesses and are lacking capabilities of the other. The Microsoft Agent Framework combines the best of both worlds with enterprise grade features, tool \u0026amp; protocol interoperability and deterministic \u0026amp; dynamic orchestration patterns. Got your attention? In this blog post, I will give the framework a try and highlight important capabilities and sources.\nOverview # The release blog of the new Microsoft Agent Framework can be found here. In short, the new framework is an open-source development kit for building multi-agent apps in .NET and Python. It is the successor of the two well-known open-source frameworks Semantic Kernel and AutoGen. It incorporates AutoGen\u0026rsquo;s powerful orchestration capabilities with Semantic Kernel\u0026rsquo;s enterprise features. Additionally, it introduces Workflows that provide control over multi-agent execution paths.\nWhat you get from the new framework:\nInteroperability - MCP, A2A, OpenAPI, cloud-agnostic runtime and more Advanced orchestration patterns - sequential, concurrent, handoff, Magentic orchestration and workflow based execution Extensibility - built in connectors, pluggable memory modules, declarative agents, community innovation and more Production ready - observability, security \u0026amp; compliance, human in the loop and more The new Framework offers capabilities in two main categories and it is important to understand the difference between them.\nAI Agents are LLM-driven entities that dynamically process user inputs, call tools and MCP servers to perform actions, and generate context-aware responses using model providers like Azure OpenAI, OpenAI, and Azure AI.\nWorkflows are graph-based, predefined sequences that connect multiple agents, human interactions, and external systems to perform complex, multi-step tasks with controlled execution paths, supporting features like type-based routing, nesting, checkpointing, and human-in-the-loop scenarios.\nFor those of you who want to migrate from Semantic Kernel or AutoGen to the new Microsoft Agent Framework, have a look at the documentation links here:\nMigrate from AutoGen Migrate from Semantic Kernel NOTE: The Microsoft Agent Framework reached general availability with the 1.0.0 release on April 3, 2026 for both Python and .NET, so you no longer need the --pre flag for the core packages. The GA announcement and the Python significant changes guide have the details.\nRequirements # To follow along, all you need is:\nPython 3.10 or later A Microsoft Foundry project (formerly Azure AI Foundry) with a model deployment The Azure CLI (az), sign in once with az login A local IDE such as VS Code In terms of infrastructure deployment, you can find Bicep templates creating the Foundry resources and model deployments here.\nInstalling the framework # The core packages plus the OpenAI, Azure OpenAI and Foundry providers are stable in 1.0.0 release, so a plain pip install works for them.\npip install agent-framework Several other providers are still in preview and need the --pre flag, for example agent-framework-anthropic (Claude), agent-framework-bedrock, agent-framework-gemini and agent-framework-ollama. The same goes for integration packages like agent-framework-copilotstudio. When in doubt, check the individual package README, since the install command there tells you whether --pre is required.\nI use the Microsoft Foundry as the provider and point agent-framework at a model I deployed in my Foundry project via two environment variables in a local .env file:\nFOUNDRY_PROJECT_ENDPOINT=\u0026#34;https://\u0026lt;your-resource\u0026gt;.services.ai.azure.com/api/projects/\u0026lt;your-project\u0026gt;\u0026#34; FOUNDRY_MODEL=\u0026#34;\u0026lt;your-model-deployment-name\u0026gt;\u0026#34; NOTE: Agent Framework does not load .env files automatically, so call load_dotenv() at the start of your script.\nCoding # For this demo, I will build a small healthcare triage app with the new framework and the Magentic orchestration pattern. This is the same scenario I use in my workshop: instead of one all-knowing agent, a manager delegates to specialist doctors (a cardiologist, a neurologist and a general practitioner) and pulls their input together into a single assessment. If you have seen my earlier Semantic Kernel multi-agent post, the idea will feel familiar.\nI packaged a full, hands-on version of everything below, including setup, code and lessons, into a workshop you can clone and run yourself:\nMicrosoft Agent Framework Workshop Let me walk through the parts that stood out to me when moving over from Semantic Kernel.\nNo more Kernel, and stateless agents # In Semantic Kernel, the Kernel was the central object you registered services and plugins on. In Agent Framework that abstraction is gone. You create a chat client, hand it to an Agent, and you are up and running:\nimport asyncio import os from dotenv import load_dotenv from agent_framework import Agent from agent_framework.foundry import FoundryChatClient from azure.identity import AzureCliCredential load_dotenv() async def main() -\u0026gt; None: # Create a FoundryChatClient to connect to the Foundry project client = FoundryChatClient( project_endpoint=os.environ[\u0026#34;FOUNDRY_PROJECT_ENDPOINT\u0026#34;], model=os.environ.get(\u0026#34;FOUNDRY_MODEL\u0026#34;), credential=AzureCliCredential(), ) # Create an agent with instructions for a healthcare assistant agent = Agent( client=client, name=\u0026#34;HealthBot\u0026#34;, instructions=\u0026#34;You are a friendly healthcare assistant. Give short, practical answers and always remind the user to consult a real doctor.\u0026#34;, ) # Execute agent run and print the result print(\u0026#34;User: What are common causes of a persistent headache?\u0026#34;) result = await agent.run(\u0026#34;What are common causes of a persistent headache?\u0026#34;) print(f\u0026#34;HealthBot: {result}\u0026#34;) if __name__ == \u0026#34;__main__\u0026#34;: asyncio.run(main()) One thing to be aware of: an Agent is stateless by default, meaning it does not remember previous turns. For multi-turn conversations you give it an AgentSession, and for automatic memory you add a history provider, both of which I cover below. Small change from Semantic Kernel, but good to know up front.\nPick your model provider # The Agent interface is the same no matter which model provider you use, you just swap the client. Handy when you want to move a prototype from OpenAI to your own Foundry deployment:\nProvider Client class Microsoft Foundry FoundryChatClient Azure OpenAI / OpenAI OpenAIChatClient / OpenAIChatCompletionClient Anthropic Claude AnthropicClient Amazon Bedrock BedrockChatClient Google Gemini GeminiChatClient Ollama (local) OllamaChatClient NOTE: In the latest Python release, Azure OpenAI uses the same agent_framework.openai clients as OpenAI, you select Azure by passing credential or azure_endpoint.\nKeep in mind that not every provider supports every feature. Things like function tools, structured outputs, a code interpreter, file search, MCP tools and background responses vary from provider to provider, so it is worth checking the Providers Overview and its provider comparison table before you commit to one.\nConversations and memory # By default every agent.run(...) call is independent, so the agent has no memory of what you said before. To hold an actual conversation, you create an AgentSession and pass it on each turn. The session carries the conversation state for you:\nimport asyncio import os from dotenv import load_dotenv from agent_framework import Agent from agent_framework.foundry import FoundryChatClient from azure.identity import AzureCliCredential load_dotenv() async def main() -\u0026gt; None: client = FoundryChatClient( project_endpoint=os.environ[\u0026#34;FOUNDRY_PROJECT_ENDPOINT\u0026#34;], model=os.environ.get(\u0026#34;FOUNDRY_MODEL\u0026#34;), credential=AzureCliCredential(), ) agent = Agent( client=client, name=\u0026#34;SymptomChecker\u0026#34;, instructions=\u0026#34;You are a symptom-checker. Remember what the patient tells you.\u0026#34;, ) # Create a session to maintain conversation state across turns session = agent.create_session() # turn 1 print(\u0026#34;User: Hi, I am Alex and I have a mild headache.\u0026#34;) result1 = await agent.run( \u0026#34;Hi, I am Alex and I have a mild headache.\u0026#34;, session=session ) print(f\u0026#34;SymptomChecker: {result1}\u0026#34;) # turn 2 print(\u0026#34;User: What did I just tell you about my symptoms?\u0026#34;) result2 = await agent.run( \u0026#34;What did I just tell you about my symptoms?\u0026#34;, session=session ) print(f\u0026#34;SymptomChecker: {result2}\u0026#34;) if __name__ == \u0026#34;__main__\u0026#34;: asyncio.run(main()) Because both calls share the same session, the second answer still knows about the headache from the first turn. Drop the session argument and that context is gone again. One session per user or per conversation also keeps their histories from mixing.\nAn AgentSession holds context for the length of a single script run. To make the agent store and reload the history automatically, add a context provider. The built-in InMemoryHistoryProvider is the simplest one, just pass it to the Agent constructor:\nfrom agent_framework import Agent, InMemoryHistoryProvider agent = Agent( client=client, name=\u0026#34;MemoryBot\u0026#34;, instructions=\u0026#34;You are a patient-intake assistant. Remember everything the patient tells you.\u0026#34;, context_providers=[ InMemoryHistoryProvider(\u0026#34;memory\u0026#34;, load_messages=True), ], ) With load_messages=True the provider reloads the full history before each LLM call, so the model always sees the whole conversation. History providers are a kind of context provider, the extension point that runs before and after every agent invocation. For production you would swap InMemoryHistoryProvider for a database-backed one (for example Cosmos DB) so the agent remembers across sessions, the interface stays the same.\nGiving agents tools with function calling # A chat-only agent is nice, but the real power shows up when the agent can do things. In Agent Framework you give an agent tools by handing it plain Python functions decorated with @tool. No plugin base classes and no hand-written JSON schema: the framework reads your type hints, docstring and the @tool metadata and builds the tool definition for you.\nimport asyncio import os from typing import Annotated from dotenv import load_dotenv from agent_framework import Agent, tool from agent_framework.foundry import FoundryChatClient from azure.identity import AzureCliCredential from pydantic import Field load_dotenv() # Define a tool that can be used by the agent to look up information about common symptoms @tool(name=\u0026#34;check_symptom\u0026#34;, description=\u0026#34;Look up general information about a common symptom.\u0026#34;) def check_symptom( symptom: Annotated[str, Field(description=\u0026#34;A symptom to look up, e.g. \u0026#39;headache\u0026#39;\u0026#34;)], ) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Return short, general information about a common symptom.\u0026#34;\u0026#34;\u0026#34; knowledge = { \u0026#34;headache\u0026#34;: \u0026#34;Often caused by tension, dehydration or lack of sleep.\u0026#34;, \u0026#34;fever\u0026#34;: \u0026#34;A common sign the body is fighting an infection.\u0026#34;, } return knowledge.get(symptom.lower(), \u0026#34;No information available for that symptom.\u0026#34;) async def main() -\u0026gt; None: # Create a FoundryChatClient to connect to the Foundry project client = FoundryChatClient( project_endpoint=os.environ[\u0026#34;FOUNDRY_PROJECT_ENDPOINT\u0026#34;], model=os.environ.get(\u0026#34;FOUNDRY_MODEL\u0026#34;), credential=AzureCliCredential(), ) # Create an agent with the symptom-checking tool agent = Agent( client=client, name=\u0026#34;HealthBot\u0026#34;, instructions=\u0026#34;You are a healthcare assistant. Use your tools when a user asks about a symptom.\u0026#34;, tools=[check_symptom], ) # Execute agent run and print the result print(\u0026#34;User: I have a headache, what could it be?\u0026#34;) result = await agent.run(\u0026#34;I have a headache, what could it be?\u0026#34;) print(f\u0026#34;Agent: {result}\u0026#34;) if __name__ == \u0026#34;__main__\u0026#34;: asyncio.run(main()) When the model decides it needs the tool, the framework calls check_symptom for you, feeds the result back into the conversation, and lets the model finish its answer. The Annotated + Field description is exactly what the model sees as the parameter documentation, so write it like a mini prompt.\nTIP: Function tools are the most common type, but they are not the only one. Depending on the provider, Agent Framework also supports web search, file search, a code interpreter and tools from MCP servers, all wired in through the same tools=[...] mechanism.\nOrchestrating a team of specialists with Magentic # Now the fun part. The Magentic pattern (inspired by AutoGen\u0026rsquo;s Magentic-One) puts a manager agent in charge: it plans the task, decides which specialists to call and how many rounds are needed, then synthesises the result. No fixed edges, no routing logic, the manager figures it out.\nflowchart LR U[Patient case] --\u003e M[Manager] M --\u003e C[Cardiologist] M --\u003e N[Neurologist] M --\u003e G[General Practitioner] C --\u003e M N --\u003e M G --\u003e M M --\u003e A[Final assessment] Here is the whole thing: three specialist doctors and a manager that coordinates them, wired together with MagenticBuilder:\nimport asyncio import os from dotenv import load_dotenv from agent_framework import Agent from agent_framework.foundry import FoundryChatClient from agent_framework.orchestrations import MagenticBuilder from azure.identity import AzureCliCredential load_dotenv() async def main() -\u0026gt; None: client = FoundryChatClient( project_endpoint=os.environ[\u0026#34;FOUNDRY_PROJECT_ENDPOINT\u0026#34;], model=os.environ.get(\u0026#34;FOUNDRY_MODEL\u0026#34;), credential=AzureCliCredential(), ) # Specialist agents cardiologist = Agent( client=client, name=\u0026#34;Cardiologist\u0026#34;, instructions=( \u0026#34;You are a cardiologist. Provide analysis only for heart-related \u0026#34; \u0026#34;symptoms. If the case is not cardiac, say so briefly.\u0026#34; ), ) neurologist = Agent( client=client, name=\u0026#34;Neurologist\u0026#34;, instructions=( \u0026#34;You are a neurologist. Provide analysis only for neurological \u0026#34; \u0026#34;symptoms. If the case is not neurological, say so briefly.\u0026#34; ), ) general_practitioner = Agent( client=client, name=\u0026#34;GeneralPractitioner\u0026#34;, instructions=( \u0026#34;You are a general practitioner. Provide a holistic assessment \u0026#34; \u0026#34;and summarise recommendations from the specialists.\u0026#34; ), ) # The manager coordinates the specialists manager = Agent( client=client, name=\u0026#34;Manager\u0026#34;, instructions=( \u0026#34;You coordinate a team of medical specialists. Delegate tasks to the \u0026#34; \u0026#34;right specialist based on the patient\u0026#39;s symptoms and synthesise their \u0026#34; \u0026#34;responses into a final assessment.\u0026#34; ), ) # Build the Magentic orchestration workflow = MagenticBuilder( participants=[cardiologist, neurologist, general_practitioner], manager_agent=manager, intermediate_output_from=[cardiologist, neurologist, general_practitioner], ).build() patient_case = ( \u0026#34;A 55-year-old patient reports chest tightness, occasional headaches, \u0026#34; \u0026#34;and numbness in the left arm that started two weeks ago.\u0026#34; ) print(f\u0026#34;Patient case: {patient_case}\\n\u0026#34;) # Stream each specialist\u0026#39;s contribution as it arrives current_agent = None async for event in workflow.run(patient_case, stream=True): if event.type in (\u0026#34;output\u0026#34;, \u0026#34;intermediate\u0026#34;) and hasattr(event.data, \u0026#34;text\u0026#34;) and event.data.text: name = getattr(event.data, \u0026#34;author_name\u0026#34;, None) if name and name != current_agent: current_agent = name print(f\u0026#34;\\n\\n[{current_agent}]\\n\u0026#34;) print(event.data.text, end=\u0026#34;\u0026#34;, flush=True) print() if __name__ == \u0026#34;__main__\u0026#34;: asyncio.run(main()) That is it. MagenticBuilder takes a list of participants and a manager_agent, and intermediate_output_from=[...] surfaces each specialist\u0026rsquo;s contribution as an intermediate event so I can stream it as the manager works through the case. If you prefer fixed pipelines instead of an AI-planned flow, the framework also ships SequentialBuilder, ConcurrentBuilder, HandoffBuilder and GroupChatBuilder, same idea, different control style.\nTesting it visually with DevUI # Running scripts in a terminal is fine, but the framework also ships DevUI, a lightweight browser interface where you can chat with and debug your agents and workflows without building a frontend first. You point it at your agent module and it gives you a proper chat window, tool-call inspection and more. It made iterating on the agents a lot quicker for me.\nOn top of that, the Foundry Toolkit extension for VS Code (formerly the AI Toolkit) adds an OpenTelemetry trace viewer, so once you call configure_otel_providers() you can watch every LLM call, tool invocation and token count for a run right inside the editor. I cover both DevUI and tracing in the workshop.\nWrap-up # The Microsoft Agent Framework really does bring the best of both worlds together: AutoGen\u0026rsquo;s orchestration muscle and Semantic Kernel\u0026rsquo;s enterprise features, now stable and shipping as 1.0.0. For me the highlights are the leaner developer experience without the Kernel, being able to swap model providers in one line, and Workflows plus Magentic orchestration for coordinating multiple agents, with DevUI and the Foundry Toolkit making it genuinely pleasant to build and debug.\nIf you want to get hands-on, clone my Microsoft Agent Framework Workshop and work through it lesson by lesson. And if you are coming from Semantic Kernel or AutoGen, the official migration guides should make the move a good bit easier, do not expect it to be completely painless, but they give you a solid starting point. This space is still moving fast, so treat this as a snapshot, but the foundation feels solid now. Happy building!\nSources # Everything I referenced while writing this post:\nIntroducing Microsoft Agent Framework (release blog) Microsoft Agent Framework 1.0 (GA announcement) Microsoft Agent Framework overview Providers overview Python 2026 significant changes guide Migrate from AutoGen Migrate from Semantic Kernel Foundry infrastructure Bicep templates Agent Framework GitHub repo Quickstart guide Foundry Toolkit for VS Code My Microsoft Agent Framework Workshop My Semantic Kernel multi-agent post And a few videos worth watching:\n","date":"14 July 2026","externalUrl":null,"permalink":"/posts/agent-framework/","section":"Posts","summary":"Intro # The Microsoft Agent Framework has been around for a while now and, with the 1.","title":"Intro to Microsoft Agent Framework","type":"posts"},{"content":"","date":"14 July 2026","externalUrl":null,"permalink":"/tags/multi-agent/","section":"Tags","summary":"","title":"Multi-Agent","type":"tags"},{"content":"","date":"14 July 2026","externalUrl":null,"permalink":"/tags/semantic-kernel/","section":"Tags","summary":"","title":"Semantic Kernel","type":"tags"},{"content":"","date":"14 July 2026","externalUrl":null,"permalink":"/tags/vs-code/","section":"Tags","summary":"","title":"VS Code","type":"tags"},{"content":" Intro # In this short blog post, I want to share with you some of the news and changes from Microsoft Build 2025 around Azure AI Foundry. For example, the Azure AI Foundry Agent Service aka Azure AI Agent Service is now GA. The blog post for the GA announcement can be found here. Additionally, there were plenty of new public preview features for Azure AI Foundry announced, such as Connected Agents, Model Router, Tracing and many more. A summary of the Azure AI Foundry announcements can be found here. I am not going to cover every feature in this blog post, instead we are going to focus on the introduced changes and tease some of the VS Code related capabilities.\nNew resources \u0026amp; changes # As mentioned above, besides all the new announcements, there were some changes around the Azure AI Foundry resources and APIs. The former construct was built with Hub-based projects hosted in an Azure AI Foundry Hub.\nThis is now being replaced by Foundry projects built on an Azure AI Foundry resource. This provides a simplified setup and coding, improved governance and management.\nCompared to the five Azure resources required for the old construct (Key Vault, AI Services, Storage Account, Azure AI Hub, Azure AI Project), it only needs one Azure AI Foundry and a project resource for the new approach.\nThere is a great blog article covering the changes, here. In terms of capabilities of the different project types, there is a comparison matrix and additional information to be found, here.\nAPI \u0026amp; SDK changes # From an API and SDK perspective, there were also a couple of changes introduced to unify the development of the core building blocks of your AI application. Here are a few examples:\nConnecting via the PROJECT_CONNECTION_STRING is not possible anymore and got changed to PROJECT_ENDPOINT, see following example: # create ai project client project_client = AIProjectClient( endpoint=PROJECT_ENDPOINT, credential=DefaultAzureCredential() ) The AI Foundry project endpoint url can be found under the Overview tab of your project, see following screenshot:\nCreating a thread changed from project_client.agents.create_thread to project_client.agents.threads.create, see following example: # create a thread thread = project_client.agents.threads.create() Creating a message changed from project_client.agents.create_message to project_client.agents.messages.create # create a message message = project_client.agents.messages.create( thread_id=thread.id, role=\u0026#34;user\u0026#34;, content=\u0026#34;Who is the greatest Basketball player of all time?\u0026#34;, ) Creating a run changed from project_client.agents.create_and_process_run to project_client.agents.runs.create_and_process, see example here: # ask the agent to perform work on the thread run = project_client.agents.runs.create_and_process(thread_id=thread.id, agent_id=agent.id) \u0026hellip; The blog post here introduces the new Azure AI Foundry API and SDK with various examples.\nIf you were following my previous blog posts about Azure AI Agent Service and Azure AI Foundry, you will notice that the code on the blog posts is still using the old approach. However, on the referenced GitHub repo, I have already changed the code to work with the latest changes.\nBuild and deploy Azure AI Foundry Agents via VS Code # If you haven\u0026rsquo;t done it already, you should take a look at the Azure AI Foundry VS Code extension (currently in Preview). It allows you to build, test, and deploy Azure AI Foundry Agents from within VS Code. Agents and tools can be defined via yaml files as you can see in the following screenshot:\nAdditionally, they can be defined and deployed via a graphical editor in VS Code, as you can see here:\nCheck out the following blog post and the documentation link here if you are keen to learn more about working with Azure AI Foundry from within VS Code.\nOpen in VS Code workflow # Another nice little new feature is the option to open a VS Code for the Web session directly from the Azure AI Foundry Agent playground.\nFrom the sample code page, you can press Open in VS Code.\nThis will open a VS Code for the Web session.\nIf you want to learn more about this feature, have a look at the following blog post.\nInfrastructure templates # In terms of infrastructure deployment, you can find bicep templates creating the new Azure AI Foundry resources here.\nSummary # If you are already using Azure AI Foundry and you are using the Hub-based construct, you can continue using it, as it will remain accessible for now. However, if you are not using any feature that is solely available on Hub-based projects, you should consider moving to the new Foundry-based projects, as new capabilities and services in GA (e.g., Foundry Agent Service \u0026amp; Foundry API) will only be made available there.\nAs I spend most of my time in VS Code, I love the integration to Azure AI Foundry via the extension. It simply is a big time saver, and I can iterate fast from within VS Code. If you haven\u0026rsquo;t, you should definitely check it out.\nSources # Announcing General Availability of Azure AI Foundry Agent Service What\u0026rsquo;s new in Azure AI Foundry - Microsoft Build 2025 Build recap: new Azure AI Foundry resource, Developer APIs and Tools Documentation \u0026amp; comparison matrix Azure AI Foundry projects Coding the Future of AI with Azure AI Foundry API and SDK Create Enterprise AI Agents with Azure AI Foundry VSCode Extension Documentation - Work with Azure AI Foundry Agent Service in Visual Studio Code (Preview) Azure AI Foundry bicep templates Documentation - VS Code for the Web Code quicker with Azure AI Foundry playgrounds and Visual Studio Code ","date":"3 July 2025","externalUrl":null,"permalink":"/posts/foundry-news/","section":"Posts","summary":"Intro # In this short blog post, I want to share with you some of the news and changes from Microsoft Build 2025 around Azure AI Foundry.","title":"Azure AI Foundry News \u0026 Changes","type":"posts"},{"content":"","date":"3 July 2025","externalUrl":null,"permalink":"/tags/ms-build/","section":"Tags","summary":"","title":"MS Build","type":"tags"},{"content":"","date":"23 April 2025","externalUrl":null,"permalink":"/tags/function-calling/","section":"Tags","summary":"","title":"Function Calling","type":"tags"},{"content":"","date":"23 April 2025","externalUrl":null,"permalink":"/tags/plugins/","section":"Tags","summary":"","title":"Plugins","type":"tags"},{"content":" Intro # These days, LLMs are not good enough anymore. We build agentic applications that perform complex tasks with up-to-date data and external systems. This is why we need multiple specialized AI Agents that can access all sorts of APIs and data sources to provide additional context to our LLMs. In this blog post, I want to take a closer look at Semantic Kernel Function Calling \u0026amp; Plugins and how it can help in such scenarios.\nSemantic Kernel # Semantic Kernel is a lightweight multiagent orchestration framework. It allows us to build enterprise-grade multiagent AI apps by combining and integrating multiple AI models and systems. In one of my last blog posts, I explained how to use Semantic Kernel to have multiple AI agents interacting with each other via a group chat. Today we are going to look at extensibility!\nFunction Calling # What is function calling? Function calling is a feature that enables AI models to request the execution of specific functions, allowing them to perform actions based on user inputs. Semantic Kernel then marshals the request to the appropriate function in your codebase and returns the results to the LLM, so the LLM can generate a final response. This capability is particularly useful for enabling AI to interact with and invoke your APIs. It is a native feature in most of the latest large language models (LLMs), facilitating planning and execution of tasks.\nPlugins # Aren\u0026rsquo;t plugins the same? Yes and no. At a high level, a plugin is a collection of functions that can be made available to AI applications and services. These functions can be orchestrated by an AI application to fulfill user requests. Behind the scenes, Semantic Kernel utilizes function calling to enable this process. Additionally, plugins support dependency injection, allowing essential services such as database connections or HTTP clients to be injected into the plugin\u0026rsquo;s constructor.\nIf you do some research about Semantic Kernel plugins and function calling, you might come across skills. What used to be skills became plugins to align with the OpenAI plugin specification. You can read more about it in the following post:\nSkills to plugins: fully embracing the OpenAI plugin spec in Semantic Kernel There are three main ways to get plugins into Semantic Kernel:\nUsing native code Using an OpenAPI specification Using an MCP Server The latter two are programming languages and platform-agnostic. However, we will use the native code option as it is the easiest to start with.\nEnough theory, let\u0026rsquo;s do some coding!\nCoding # For this demonstration, I want to write a simple Semantic Kernel AI app that uses a native code plugin. For those who know me, I am a huge 🏀 Basketball fan. In my spare time, I coach a children\u0026rsquo;s team in my hometown, and I still play Basketball myself. The Basketball EuroLeague is currently in playoff mode, and I like to follow the latest game results. That\u0026rsquo;s why I am going to write a plugin that uses the EuroLeague API to provide the latest game results to my LLM.\nWriting our Semantic Kernel plugin # When writing the plugin, it is important to give special attention to function naming. It is recommended to use snake case for function names and properties, as most models are trained with Python for function calling. The model needs to understand its intent and parameters. If needed, you can, and in some cases, you should add a description to it, but it will increase the token usage. In our case, it is a simple function with only one input parameter and a self-explanatory name, get_latest_euroleague_game_results. Nevertheless, I have annotated the season input parameter for the model to know what is expected.\nAll we really need to do is create a class and annotate its methods with the @kernel_function attribute. This way, Semantic Kernel knows that this is a function that can be called by our LLM. Note that the helper function xml_to_dict is not annotated and won\u0026rsquo;t be sent to the model.\nimport xml.etree.ElementTree as ET import httpx from typing import Any, Dict, Annotated from datetime import datetime from semantic_kernel.functions import kernel_function # This plugin fetches EuroLeague game results and processes them. class EuroleaguePlugin: # This function converts an XML element and its children to a dictionary. def xml_to_dict(self, element: ET.Element) -\u0026gt; Any: \u0026#34;\u0026#34;\u0026#34;Recursively converts an XML element and its children to a dictionary.\u0026#34;\u0026#34;\u0026#34; node = {} if element.attrib: node.update(element.attrib) children = list(element) if children: for child in children: child_dict = self.xml_to_dict(child) if child.tag not in node: node[child.tag] = [] node[child.tag].append(child_dict) else: node = element.text or \u0026#34;\u0026#34; return node # This function fetches the latest six EuroLeague game results for a given season. @kernel_function async def get_latest_euroleague_game_results(self, season: Annotated[str, \u0026#34;The year of the season, e.g. 2024 for season 2024/25\u0026#34;]) -\u0026gt; Dict[str, Any]: \u0026#34;\u0026#34;\u0026#34;Fetches EuroLeague game results and returns only the last 6 games by date.\u0026#34;\u0026#34;\u0026#34; url = \u0026#34;https://api-live.euroleague.net/v1/results\u0026#34; params = {\u0026#34;season_code\u0026#34;: \u0026#34;E\u0026#34; + season} headers = {\u0026#34;Accept\u0026#34;: \u0026#34;application/xml\u0026#34;} try: async with httpx.AsyncClient() as client: response = await client.get(url, params=params, headers=headers) response.raise_for_status() root = ET.fromstring(response.text) data = self.xml_to_dict(root) # Extract the games from the parsed XML data games = data.get(\u0026#34;game\u0026#34;, []) # Parse and sort games by date def parse_date(game): # The date is a list, so get the first element date_val = game.get(\u0026#34;date\u0026#34;, [\u0026#34;\u0026#34;]) if isinstance(date_val, list): date_str = date_val[0] else: date_str = date_val try: return datetime.strptime(date_str, \u0026#34;%b %d, %Y\u0026#34;) except Exception: return datetime.min # Sort the games by date in descending order games_sorted = sorted(games, key=parse_date, reverse=True) last_6_games = games_sorted[:6] # Replace the games in the data dict with only the last 6 data[\u0026#34;game\u0026#34;] = last_6_games return data # Exception handling except Exception as e: print(f\u0026#34;Exception when making direct request: {e}\u0026#34;) return {} So far, so good. We have a plugin. What is doing exactly? It is getting the latest six Euroleague game results of a specific season. The helper function xml_to_dict converts the XML response from the EuroLeague API to a Python dictionary, so it is easier to process for the model. We then sort the data in descending order by date and truncate it, leaving only the last six games.\nWriting our Semantic Kernel AI app # Now that we have a plugin, let\u0026rsquo;s write our AI app that will use the plugin. This script uses Azure OpenAI chat completion, adds our Euroleague plugin to the kernel, and creates a chat history to interact with the assistant.\nAdditionally, we define the AzureChatPromptExecutionSettings to configure how prompts are executed when interacting with Azure OpenAI. It allows you to set options such as function calling behavior, temperature, and other model parameters. In our case, we set the FunctionChoiceBehavior to auto. This setting lets the model automatically select the most relevant function based on the user\u0026rsquo;s input.\nLastly, the script sends a user message requesting the latest Euroleague game results and prints the AI\u0026rsquo;s response.\nimport asyncio import os import euroleague_plugin from dotenv import load_dotenv from semantic_kernel import Kernel from semantic_kernel.connectors.ai.open_ai import AzureChatCompletion, AzureChatPromptExecutionSettings from semantic_kernel.connectors.ai import FunctionChoiceBehavior from semantic_kernel.contents import ChatHistory load_dotenv() # Azure OpenAI config (set your environment variables) AZURE_OPENAI_KEY = os.getenv(\u0026#34;AZURE_OPENAI_KEY\u0026#34;) AZURE_OPENAI_ENDPOINT = os.getenv(\u0026#34;AZURE_OPENAI_ENDPOINT\u0026#34;) AZURE_OPENAI_DEPLOYMENT = os.getenv(\u0026#34;AZURE_OPENAI_DEPLOYMENT\u0026#34;) async def main(): # Initialize the kernel kernel = Kernel() # Add Azure OpenAI chat completion chat_completion = AzureChatCompletion( deployment_name=AZURE_OPENAI_DEPLOYMENT, api_key=AZURE_OPENAI_KEY, endpoint=AZURE_OPENAI_ENDPOINT, ) kernel.add_service(chat_completion) # Add a plugin (the EuroleaguePlugin class is defined above) kernel.add_plugin( euroleague_plugin.EuroleaguePlugin(), plugin_name=\u0026#34;Euroleague\u0026#34; ) # Enable planning execution_settings = AzureChatPromptExecutionSettings() execution_settings.function_choice_behavior = FunctionChoiceBehavior.Auto() # Create a history of the conversation history = ChatHistory() history.add_user_message(\u0026#34;Please give me the latest Euroleague game results for the 2024 season.\u0026#34;) # Get the response from the AI result = await chat_completion.get_chat_message_content( chat_history=history, settings=execution_settings, kernel=kernel, ) # Print the results print(\u0026#34;Assistant \u0026gt; \u0026#34; + str(result)) # Add the message from the agent to the chat history history.add_message(result) if __name__ == \u0026#34;__main__\u0026#34;: asyncio.run(main()) Let\u0026rsquo;s see if our code works as expected and if our AI app will indeed use the plugin to fetch the latest game results from the EuroLeague API.\nRunning our Semantic Kernel AI app # Note: We haven\u0026rsquo;t covered how to create the Azure OpenAI Service or deployment. If you need some guidance, you can follow the documentation here. Make sure you have an up-to-date model (e.g., gpt-4o) deployed and a local .env file with the endpoint, key and deployment name filled out.\nI will simply execute the app in my terminal and see if we get the expected response back.\n➜ python app3.py Assistant \u0026gt; Here are the latest Euroleague game results for the 2024 season: 1. **April 24, 2025** - **Panathinaikos Aktor Athens** vs. **Anadolu Efes Istanbul**: 76 - 79 - **Fenerbahce Beko Istanbul** vs. **Paris Basketball**: 89 - 72 2. **April 23, 2025** - **Olympiacos Piraeus** vs. **Real Madrid**: 84 - 72 - **AS Monaco** vs. **FC Barcelona**: 97 - 80 3. **April 22, 2025** - **Fenerbahce Beko Istanbul** vs. **Paris Basketball**: 83 - 78 - **Panathinaikos Aktor Athens** vs. **Anadolu Efes Istanbul**: 87 - 83 These games were part of the playoffs. Great! This is accurate and everything I wanted. The model provides me with the latest six-game results that happened just a couple of days ago. We can verify if the results are correct by checking the EuroLeague Game Center page.\nIf you want to check out the code, you can find it on my GitHub repo.\nSummary # This was a very straightforward example of how to use Semantic Kernel plugins to reach out to external systems to provide up-to-date data to our LLM. Encapsulating native code within a plugin is a simple method to equip an AI agent with capabilities that it doesn\u0026rsquo;t inherently support. This approach enables you to utilize your existing app development skills and code to enhance the functionality of your AI agents. Besides native code, MCP, and OpenAPI plugins, we can also make use of Retrieval Augmented Generation (RAG) patterns when it comes to data access and grounding. For example, we can build a data retrieval plugin using Azure AI Search as described here.\nPlenty of possibilities to make our AI agents smarter. I am especially excited about the MCP option, as I recently wrote an intro blog post to MCP, and I can\u0026rsquo;t wait to combine it with Semantic Kernel.\nSources # Semantic Kernel function calling Semantic Kernel plugins Skills to plugins: fully embracing the OpenAI plugin spec in Semantic Kernel Semantic Kernel plugins - native code Semantic Kernel plugins - OpenAPI specification Semantic Kernel plugins - MCP Server Snake case Semantic Kernel KernelFunctions Semantic Kernel ChatCompletion Semantic Kernel ChatHistory Semantic Kernel AzureChatPromptExecutionSettings Semantic Kernel FunctionChoiceBehavior Deploy Azure OpenAI via Bicep Semantic Kernel RAG plugin EuroLeague Game Center EuroLeague API ","date":"23 April 2025","externalUrl":null,"permalink":"/posts/skfc/","section":"Posts","summary":"Intro # These days, LLMs are not good enough anymore.","title":"Semantic Kernel Function Calling \u0026 Plugins: Give AI Agents Real-World Tools","type":"posts"},{"content":"","date":"11 April 2025","externalUrl":null,"permalink":"/tags/copilot/","section":"Tags","summary":"","title":"Copilot","type":"tags"},{"content":" Intro # The Model Context Protocol (MCP) is trending! What is it? Let\u0026rsquo;s check it out. MCP is an open-source project launched in November 2024. It defines an open standard for AI applications to connect to their tools and data. I heard someone referring to it as the USB adapter for AI systems because it provides a very much needed standardization. Making your AI models smarter typically requires custom code, which can be challenging when it comes to scaling and maintaining the integrations. Most likely, your teams will end up writing the same integration multiple times for different AI use-cases. MCP allows for a \u0026ldquo;write once, work everywhere\u0026rdquo; approach, which will save development efforts.\nAdditionally, I recommend watching the John Savill\u0026rsquo;s \u0026ldquo;Model Context Protocol Overview - Why You Care!\u0026rdquo; video:\nOverview # In this blog post, I want to explain and demonstrate MCP with a simple but useful example. As everyone is currently talking about MCP and the project gains popularity, we can find more and more content about it online. For example, the collection of MCP Servers is growing by the minute. Just a couple of days ago, the public preview of Azure MCP Server was announced. As I spent most of my time in VS Code working with GitHub Copilot, I want to run the Azure MCP Server locally and use VS Code as the client that connects to it. But basics first!\nMCP Server # What is an MCP Server? MCP uses a client-server architecture. The MCP Server is where you build your connection to your APIs and resources. This is where the heavy lifting is done as you still have to write the integration. Nevertheless, the MCP Server then exposes your services as MCP Tools via the standardized Model Context Protocol and makes them available to your AI apps.\nThe MCP Servers can run local (stdio) or remote via HTTP+SSE(Server-Sent Events) transport layer. However, the remote implementation is still in early stages and is evolving fast. The recent specs already replacing HTTP+SSE with Streamable HTTP. I found a great blog post from Christian Posta explaining the change in more detail.\nUnderstanding MCP Recent Change Around HTTP+SSE If you want to host a remote MCP Server on Azure, there are already multiple blog posts and repos available that you can follow:\nHost remote MCP Servers on Azure App Service Host remote MCP Servers on Azure Container Apps Host remote MCP Servers on Azure Functions When we talk remote MCP Servers, authentication becomes very important, and I can recommend following Den Delimarsky\u0026rsquo;s blog posts and his Entra ID integration ideas. But be aware, this is also evolving fast and might be irrelevant next week.\nUsing Microsoft Entra ID To Authenticate With MCP Servers Via Sessions A list of example MCP Servers can be found here, and a quickstart on how to write your own server here. In this blog post, we are using a local MCP Server for the ease of use.\nMCP Host / Client # The MCP Host can be your AI app or LLM based coding client. It is responsible for generating tasks or queries, but does not directly interact with data sources or APIs. The MCP Host can create and manage multiple client instances.\nThe MCP Client acts as an intermediary between the MCP Host and the MCP Server. It manages connections, handles communication, discovers and executes tools, and facilitates resource access. This ensures that the AI model can perform its tasks efficiently and effectively by leveraging the capabilities of the MCP Servers.\nIn summary, the MCP Client acts as a bridge between the AI model and the MCP Servers.\nflowchart LR subgraph \"MCP Host\" direction LR A[MCP Client] B[MCP Server] C[MCP Client] D[MCP Server] G[MCP Client] end E[Resource] F[API] I[API] H[Remote MCP Server] A--\u003eB B\u003c--\u003eE C--\u003eD D\u003c--\u003eF G--\u003eH H\u003c--\u003eI A list of applications that support MCP can be found here and a quickstart on how to build your own client here. Additionally, you can find more about the MCP architecture in the documentation and from the latest specifications as of writing this blog post.\nMCP Capabilities # What are the capabilities that can be exposed by MCP Servers and used by MCP Clients?\nMCP Tools: Functions that can be invoked by your LLMs. This allows for automation and extensibility into the outside world. MCP Resources: Data that can be accessed by your LLMs. This allows for providing data such as files, databases, log files etc\u0026hellip; as context to your LLMs. MCP Prompts: Reusable prompt templates that can be used by your LLMs. This allows to standardize the LLM interactions. MCP Sampling: Allows the MCP Server to request completions from the client. This is very early stages and the client support is still lacking. However, this can become very powerful as it enables bidirectional communication. Implementation # Now that we have covered the basics, let\u0026rsquo;s start with the implementation. As mentioned above, I want to use the Azure MCP Server with VS Code and GitHub Copilot.\nPrerequisites # If you want to follow along, make sure you have the following prerequisites in place:\nAzure Subscription Azure CLI GitHub account and a GitHub Copilot free plan or higher VS Code with GitHub Copilot and Agent Mode enabled Make sure VS Code GitHub Copilot is signed in with your GitHub account, agent mode is enabled and selected on the bottom of the chat window.\nConfiguration # First, open a terminal in VS Code and sign in to our Azure Subscription with the az logincommand. Create an empty folder and create a .vscode folder inside the first folder. Create a mcp.json file within the .vscode folder and add the following content to it. { \u0026#34;servers\u0026#34;: { \u0026#34;Azure MCP Server\u0026#34;: { \u0026#34;command\u0026#34;: \u0026#34;npx\u0026#34;, \u0026#34;args\u0026#34;: [ \u0026#34;-y\u0026#34;, \u0026#34;@azure/mcp\u0026#34;, \u0026#34;server\u0026#34;, \u0026#34;start\u0026#34; ] } } } Note: I experienced some processor architecture related issues when using \u0026ldquo;@azure/mcp@latest\u0026rdquo; package as suggested by the blog post or the repo. I am running on a MacBook with M2 processor, and it was not finding the mcp-darwin-arm64 package. Nevertheless, after removing the @latest it was working, and the version is also identical (0.0.10 at the time of writing this blog post).\nAfter saving the mcp.json you should see a little \u0026ldquo;Start\u0026rdquo; button appearing over \u0026ldquo;Azure MCP Server\u0026rdquo;. Press it! This will start the Azure MCP Server, and it will discover the MCP Tools provided by it. Now we should see the discovered tools in our Agent mode chat window. We can inspect the available tools by clicking on the tools icon. So far so good, let\u0026rsquo;s try the MCP Tools.\nExecution # Now that we have everything in place, let\u0026rsquo;s ask our Copilot Agent something about Azure, e.g. \u0026ldquo;list all my Azure Resource Groups\u0026rdquo;. It will automatically realize that it has MCP Tools available that can help here. The Agent will also recognize that it first has to get the Azure Subscription by running the \u0026ldquo;azmcp-subscription-list\u0026rdquo; tool. We can get more details about the tool by clicking on \u0026ldquo;\u0026gt; Run\u0026rdquo;.\nAfter approving the tool execution by clicking \u0026ldquo;Continue\u0026rdquo; it will run the first tool and ask to execute the second tool to list my Resource Groups. As soon as we approve the next step, it will show us a list with our Resource Groups.\nAwesome, we can now use the MCP Tools provided by our Azure MCP Server to execute azd commands directly or query logs and more. Combine this with the Azure extension for GitHub Copilot and you have a powerful Azure toolset directly in VS Code GitHub Copilot.\nSummary # To summarize, MCP provides a very much needed standardization in the ever-growing AI Agent space. We can already see that multiple vendors adopt it and provide MCP Servers for their solutions. In addition, the amount of community provided MCP Servers is growing by the minute. In an ideal case, you don\u0026rsquo;t have to write your own MCP Server, you just look it up in an MCP Server registry.\nHere is a list of popular MCP Server registries:\nhttps://mcp.so/servers https://glama.ai/mcp/servers ​ https://smithery.ai/ https://www.pulsemcp.com/servers If you wonder, is it is possible to use Azure AI Agent Service with MCP, it is. There is a great blog post available here. It describes how to use Azure AI Agents as MCP Tools. Which might sounds strange in the first moment as it works the other way around. In this scenario, the MCP Server uses the agent or multiple agents as tools instead of the agent querying the MCP Server for exposed tools. Which makes a lot of sense if you think about the many ootb tools and Azure integrations that are available with the Azure AI Agent Service.\nBesides MCP, there is another standard that was just recently announced by Google. The A2A protocol, which caters more to the agent to agent collaboration across platforms. It is definitely worth to follow both projects and see how they evolve. There is no doubt the future of agentic apps looks bright.\nResources # Azure MCP Server MCP architecture documentation MCP Server quickstart MCP Server examples MCP Client quickstart MCP Client examples MCP specifications 26.03.2025 MCP Tools MCP Resources MCP Prompts MCP Sampling Understanding MCP Recent Change Around HTTP+SSE Host remote MCP Servers on Azure App Service Host remote MCP Servers on Azure Container Apps Host remote MCP Servers on Azure Functions Using Microsoft Entra ID To Authenticate With MCP Servers Via Sessions Azure Subscription Azure CLI documentation GitHub account VS Code GitHub Copilot GitHub Copilot Agent Mode Azure Developer CLI GitHub Copilot Azure extension Azure AI Agent Service MCP with Azure AI Agent Service A2A protocol ","date":"11 April 2025","externalUrl":null,"permalink":"/posts/mcp/","section":"Posts","summary":"Intro # The Model Context Protocol (MCP) is trending!","title":"Model Context Protocol: The USB Adapter for AI Apps and Agents","type":"posts"},{"content":" Intro # In my last post, I have covered Azure AI Agent Service and how it can be used to easily build and run AI agents on Azure. This time, we are going to look at Semantic Kernel as a framework to build, orchestrate and deploy AI agents or multi-agent applications. Semantic Kernel is an open-source development kit by Microsoft that offers a unified framework with a plugin-based architecture for easier integration and reduced complexity. It serves as efficient middleware, enabling fast development of enterprise-grade solutions by combining prompts with existing APIs.\nThe Semantic Kernel SDK is available for C#, Python and Java. More details can be found on the official GitHub repository and the official Microsoft documentation.\nIn this blog post, we are going to combine the Azure AI Agent Service with Semantic Kernel to build a multi-agent AI application with group chat functionality.\nBackground # If you are new to the topic, it can be a bit confusing as there are many options to choose from. Which framework should I use? AutoGen or Semantic Kernel? Which API should I use the Chat Completions API, the Assistants API or the Azure AI Agent Service? Well, famous last words \u0026ldquo;it depends\u0026rdquo; and \u0026ldquo;things are evolving fast\u0026rdquo;.\nThe framework discussion is manly driven by whether you need enterprise-grade support or not. If that is a yes, you should look towards Semantic Kernel. If you are still in the ideation/testing phase, and you need the latest and greatest functionality, take a look at AutoGen. Both teams are working on strategic convergence and integrations between both frameworks, as you can read in the following blog posts:\nMicrosoft’s Agentic AI Frameworks: AutoGen and Semantic Kernel Semantic Kernel Roadmap H1 2025: Accelerating Agents, Processes, and Integration AutoGen and Semantic Kernel, Part 2 In terms of API, the Chat Completions API is lightweight and stateless and can be a good fit for simple tasks. The Assistants API is stateful (managing conversation history) and can be a good fit for more complex scenarios. Azure AI Agent Service delivers all the functionality of the Assistants API plus flexible model choice, out of the box tools, tracing and more.\nRequirements # As an execution engine, we will use the Azure AI Agent Service. We are not going to cover the infrastructure requirements in this blog post. Nevertheless, if you want to get started quickly, simply deploy this bicep template for an standard Azure AI Agent deployment.\nPrepare local dev environment # For our local development environment, we need to install the following packages:\npip install python-dotenv azure-identity semantic-kernel[azure] Next, we will use a local .env file to specify some variables to connect to our Azure AI Foundry Project.\nAZURE_AI_AGENT_PROJECT_CONNECTION_STRING=\u0026#34;your_project_connection_string\u0026#34; AZURE_AI_AGENT_MODEL_DEPLOYMENT_NAME=\u0026#34;your_model_deployment\u0026#34; Coding # My goal is to create multiple AI agents that act as Basketball coaches. A Head Coach and an Assistant Coach that will exchange ideas and come up with a game plan for a specific game situation.\nFor this blog, we will keep it very simple and don\u0026rsquo;t play too much with plugins or extensions. We just want to create a group chat with specialized agents.\nBuilding the mulit-agent app # As always we need to add some references first. This time we are going to import the asyncio model as we need to execute the main function asynchronously. This allows the program to perform non-blocking operations, such as interacting with Azure AI services, creating agents, and managing the group chat. Additionally, we are going to import some components from the Semantic Kernel library. More about these Semantic Kernel classes can be found later in the text.\n# add references import asyncio from dotenv import load_dotenv from azure.identity.aio import DefaultAzureCredential from semantic_kernel.agents import AzureAIAgent, AzureAIAgentSettings, AgentGroupChat from semantic_kernel.agents.strategies import TerminationStrategy, SequentialSelectionStrategy from semantic_kernel.contents.utils.author_role import AuthorRole Now we are going to load the environment variables from the .env file and define the agent names and instructions as well as the task they should work on.\n# get configuration settings load_dotenv() # agent instructions HEAD_COACH = \u0026#34;HeadCoach\u0026#34; HEAD_COACH_INSTRUCTIONS = \u0026#34;\u0026#34;\u0026#34; You are a Basketball Head Coach that knows everything about offensive plays and strategies. You respond to specific game situations with advice on how to change the game plan. You can ask for more information about the game situation if needed. For offensive plays and strategies, you will decide on the strategy yourself. If the game situation demands a change for defensive plays and strategies, you will ask your Assistant Coach for advice. You will use the advice given by the Assistant Coach regarding defensive adjustments and your own decision for offensive adjustments to create the final game plan. RULES: - Use the instructions provided. - Prepend your response with this text: \u0026#34;head_coach \u0026gt; \u0026#34; - Do not directly answer the question if it is related to defensive strategies. Instead, ask your Assistant Coach for advice. - Do not use the words \u0026#34;final game plan\u0026#34; unless you have created a final game plan according to the instructions. - Add \u0026#34;final game plan\u0026#34; to the end of your response if you have created a final game plan according to the instructions. \u0026#34;\u0026#34;\u0026#34; ASSISTANT_COACH = \u0026#34;AssistantCoach\u0026#34; ASSISTANT_COACH_INSTRUCTIONS = \u0026#34;\u0026#34;\u0026#34; You are a Basketball Assistant Coach that knows defensive plays and strategies. You give advice to your Head Coach for specific game situations that require defensive adjustment. RULES: - Use the instructions provided. - Prepend your response with this text: \u0026#34;assistant_coach \u0026gt; \u0026#34; - You are not allowed to give advice on offensive plays and strategies. - You don\u0026#39;t decide the final game strategy and plan, you only give advice to the Head Coach. - Your advice should be clear and concise and should not include any unnecessary information. \u0026#34;\u0026#34;\u0026#34; # agent task TASK = \u0026#34;Could you please give me advice on how to change the game strategy for the next quarter? We are playing zone defense, and the other team just scored 10 points in a row. We need to change our strategy to stop them. What should we do?\u0026#34; So far, so good. We will now start with our main function. We retrieve the configuration settings with the AzureAIAgentSettings.create method and use the DefaultAzureCredential class to authenticate against our Azure Services, and lastly, we create a client for interacting with the Azure AI agent service.\nasync def main(): ai_agent_settings = AzureAIAgentSettings.create() async with ( DefaultAzureCredential(exclude_environment_credential=True, exclude_managed_identity_credential=True) as creds, AzureAIAgent.create_client(credential=creds) as client, ): The next step is to create our agents on the Azure AI Agent Service and wrap them into Semantic Kernel agents using the AzureAIAgent class.\n# create the head-coach agent on the Azure AI agent service headcoach_agent_definition = await client.agents.create_agent( model=ai_agent_settings.model_deployment_name, name=HEAD_COACH, instructions=HEAD_COACH_INSTRUCTIONS, ) # create a Semantic Kernel agent for the Azure AI head-coach agen agent_headcoach = AzureAIAgent( client=client, definition=headcoach_agent_definition, ) # create the assistant coach agent on the Azure AI agent service assistantcoach_agent_definition = await client.agents.create_agent( model=ai_agent_settings.model_deployment_name, name=ASSISTANT_COACH, instructions=ASSISTANT_COACH_INSTRUCTIONS, ) # create a Semantic Kernel agent for the assistant coach Azure AI agent agent_assistantcoach = AzureAIAgent( client=client, definition=assistantcoach_agent_definition, ) This is where the fun part begins. We are initializing a group chat via the AgentGroupChat class and adding our agents to it. We need to define who comes next and when the group chat should end. To do so, we define a termination and selection strategy. Additionally, we define which agent is contributing to the termination strategy. In our case, only the Head Coach is in charge. Furthermore, we are defining a maximum of 4 iterations until the chat will be terminated.\n# add the agents to a group chat with a custom termination and selection strategy chat = AgentGroupChat( agents=[agent_headcoach, agent_assistantcoach], termination_strategy=ApprovalTerminationStrategy( agents=[agent_headcoach], maximum_iterations=4, automatic_reset=True ), selection_strategy=SelectionStrategy(agents=[agent_headcoach,agent_assistantcoach]), ) In this section, we handle the execution of the group chat, including adding the task, invoking the chat, and performing cleanup operations. The AuthorRole.USER constant is used to explicitly identify the role of the message sender as the user to ensure clarity in the conversation flow.\ntry: # add the task as a message to the group chat await chat.add_chat_message(message=TASK) print(f\u0026#34;# {AuthorRole.USER}: \u0026#39;{TASK}\u0026#39;\u0026#34;) # invoke the chat async for content in chat.invoke(): print(f\u0026#34;# {content.role} - {content.name or \u0026#39;*\u0026#39;}: \u0026#39;{content.content}\u0026#39;\u0026#34;) finally: # cleanup and delete the agents print(\u0026#34;--chat ended--\u0026#34;) await chat.reset() await client.agents.delete_agent(agent_headcoach.id) await client.agents.delete_agent(agent_assistantcoach.id) Almost at the end, we just need to add two classes for the termination and selection function that we have used in the group chat definition. As we defined in the Head Coach agent instructions, as soon as the final game plan is ready, it should add \u0026ldquo;final game plan\u0026rdquo; to its message. We are checking if the last message in the history contains this phrase, and we will terminate the group chat.\n# class of termination strategy class ApprovalTerminationStrategy(TerminationStrategy): \u0026#34;\u0026#34;\u0026#34;A strategy for determining when an agent should terminate.\u0026#34;\u0026#34;\u0026#34; async def should_agent_terminate(self, agent, history): \u0026#34;\u0026#34;\u0026#34;Check if the agent should terminate.\u0026#34;\u0026#34;\u0026#34; return \u0026#34;final game plan\u0026#34; in history[-1].content.lower() The second class adds a selection function that defines which agent should take the next turn in the chat. If the last message comes from the User or the Assistant Coach, it is the Head Coaches turn.\n# class for selection strategy class SelectionStrategy(SequentialSelectionStrategy): \u0026#34;\u0026#34;\u0026#34;A strategy for determining which agent should take the next turn in the chat.\u0026#34;\u0026#34;\u0026#34; # select the next agent that should take the next turn in the chat async def select_agent(self, agents, history): \u0026#34;\u0026#34;\u0026#34;\u0026#34;Check which agent should take the next turn in the chat.\u0026#34;\u0026#34;\u0026#34; # the Head Coach should go after the User or the Assistant Coach if (history[-1].name == ASSISTANT_COACH or history[-1].role == AuthorRole.USER): agent_name = HEAD_COACH return next((agent for agent in agents if agent.name == agent_name), None) # otherwise it is the Assistant Coach\u0026#39;s turn return next((agent for agent in agents if agent.name == ASSISTANT_COACH), None) Lastly, we define the entry point for our app, and that it is executed as an asynchronous coroutine.\nif __name__ == \u0026#34;__main__\u0026#34;: asyncio.run(main()) Running the mulit-agent app # Ok, we should have something to play and test with. The finale code can be found on my GitHub repo here.\nLet\u0026rsquo;s see if we can get a proper game plan from our coaching staff.\n➜ python app.py # AuthorRole.USER: \u0026#39;Could you please give me advice on how to change the game strategy for the next quarter? We are playing zone defense and the other team just scored 10 points in a row. We need to change our strategy to stop them. What should we do?\u0026#39; # AuthorRole.ASSISTANT - HeadCoach: \u0026#39;head_coach \u0026gt; I\u0026#39;ll need to consult with the Assistant Coach about defensive adjustments since that\u0026#39;s not my area of expertise. Assistant Coach, what adjustments do you recommend for our zone defense to stop the opposing team who has just scored 10 points in a row? In terms of our offensive strategy, we should focus on enhancing our ball movement and executing quick passes to exploit the gaps in their defense. Let\u0026#39;s emphasize perimeter shooting and look for opportunities to drive to the basket, ensuring we spread the floor to create space. Please provide your defensive advice, and I\u0026#39;ll integrate that with our offensive strategy for the necessary adjustments.\u0026#39; # AuthorRole.ASSISTANT - AssistantCoach: \u0026#39;assistant_coach \u0026gt; Consider switching to a man-to-man defense to apply more pressure on their shooters and disrupt their rhythm. This will help limit their easy scoring opportunities and force them into more contested shots. Ensure our players communicate effectively and switch on screens. Additionally, encourage tighter closeouts on shooters to contest their shots and deny open looks. If they continue to score, we could also implement a trap to force turnovers and get out in transition.\u0026#39; # AuthorRole.ASSISTANT - HeadCoach: \u0026#39;head_coach \u0026gt; Thank you, Assistant Coach. Based on your advice, we\u0026#39;ll switch to a man-to-man defense to apply pressure and limit their scoring opportunities. We\u0026#39;ll focus on strong communication and switching on screens, as well as tighter closeouts on shooters. On the offensive side, we\u0026#39;ll continue to enhance our ball movement, emphasizing quick passes and prioritizing perimeter shooting, alongside drive opportunities. This blend of a more aggressive defensive approach and a fluid offensive strategy should help us regain control of the game. Now, let\u0026#39;s put this all together: we\u0026#39;ll implement a man-to-man defense while enhancing our offensive ball movement and exploiting gaps in their defense. final game plan\u0026#39; --chat ended-- Nice! Thanks Coaching staff! That indeed sounds like a plan to win the game in the end.\nIn Azure AI Foundry, we can see that the corresponding Azure AI Agents are getting created during the runtime and cleaned up afterward.\nSummary # This was a very simple example, but it shows how multiple agents can have different expertise and exchange ideas or knowledge about a specific topic via the Semantic Kernel group chat. Additionally, we can facilitate a structured conversation flow that allows the agents to efficiently collaborate and work on user provided tasks. Imagine that these agents would have access to different tools or knowledge sources to make them specialists for a specific task. We already looked at how to add tools (Code Interpreter Tool) to Azure AI Agents in my last blog post. The same approach can be used in combination with Semantic Kernel to make our agents even smarter.\nSources # Semantic Kernel Microsoft documentation Semantic Kernel GitHub repository Microsoft’s Agentic AI Frameworks: AutoGen and Semantic Kernel Semantic Kernel Roadmap H1 2025: Accelerating Agents, Processes, and Integration AutoGen and Semantic Kernel, Part 2 Azure AI Agent standard setup bicep template Azure AI Agent Service: Build and Run AI Agents on Azure Microsoft Learn AI Agent Fundamentals Introducing enterprise multi-agent support in Semantic Kernel Semantic Kernel Agents are now Generally Available ","date":"10 April 2025","externalUrl":null,"permalink":"/posts/multi-agent/","section":"Posts","summary":"Intro # In my last post, I have covered Azure AI Agent Service and how it can be used to easily build and run AI agents on Azure.","title":"Intro to Semantic Kernel and Multi-Agent AI Apps","type":"posts"},{"content":" Intro # AI Agents are in talks, and we are going to take a look at the Azure AI Agent Service. Why should you care? Because AI Agents are the next evolution of AI-driven applications, and they will help you to automate and execute more complicated multistep tasks completely autonomously or with a human in a loop. Here are two blog posts that are worth reading to set the scene:\nAI agents — what they are, and how they’ll change the way we work Introducing Azure AI Agent Service In this blog post, I will give you a short overview about what the Azure AI Agent Service is, what it can do for you, and how to quickly get started via the provided Python Azure AI Foundry SDK.\nNOTE: At the time of writing, the Azure AI Agent Service is in public preview.\nOverview # The Azure Azure AI Agent Service is part of Azure AI Foundry. Azure AI Foundry is a unified AI platform that allows you to manage the whole lifecycle of your AI application. Key features, are:\nRich model catalog (OpenAI, DeepSeek, Cohere, Meta, Mistral\u0026hellip;) Deploy and experiment with different models in playgrounds Seamless Integration to other Azure services such as Azure OpenAI, Azure AI Services, and Azure AI Search Project Management with features for project creation, resource management, and access control Simplified coding experience with unified SDK Build and deploy AI agents with the Azure AI Agent Service and more\u0026hellip; With the Azure AI Agent Service, Developers can easily build extensible AI Agents using out of the box tooling and integrations into Azure services. Some of the highlights are:\nFlexible model selection (not solely OpenAI models) Knowledge tools such as Azure AI Search, Grounding with Bing Search, Microsoft Fabric and file uploads Action tools such as OpenAPI 3.0 specified tools and Code Interpreter, Azure Functions and custom functions Conversation state management (providing consistent context) and more\u0026hellip; Requirements # To start with Azure AI Agent Service, we first need to deploy the necessary Azure Services and create an Azure AI Foundry Hub and Project. Additionally, we will prepare our local development environment.\nPrepare infrastructure # In this blog post, we want to focus on how to use these services from a coding perspective. Hence, we are not going to cover the infrastructure deployment in much detail. However, we will simply use the provided bicep template from the Azure-Samples repository here.\nAfter a successful deployment, we should see the following services in your resource group.\nThe Azure AI Foundry portal can be found under https://ai.azure.com/. We should now have an Azure AI Foundry Hub and Project created for us.\nIf you want to learn more about Azure AI Foundry Hubs and Projects, visit the documentation page here.\nWe can also view the connected resources that have been created as part of the deployment and can be used for the AI agents that we want to build.\nPrepare local dev environment # For our local development environment, we need to install the following packages:\npip install python-dotenv azure-ai-projects azure-identity Next, we will use a local .env file to specify some variables to connect to our Azure AI Foundry Project.\nAZURE_AI_AGENT_PROJECT_CONNECTION_STRING=\u0026#34;your_project_connection_string\u0026#34; AZURE_AI_AGENT_MODEL_DEPLOYMENT_NAME=\u0026#34;your_model_deployment\u0026#34; Coding # If everything is running we can start coding. My goal is to build an AI Agent that acts as a Basketball Assistant Coach. For those who don\u0026rsquo;t know me, I am a big Basketball fan and combing two things that I love is pure excitement for me.\nBuilding the app # First we need to import the necessary classes from the SDK. In this case we are using the python SDK.\nimport os from dotenv import load_dotenv from azure.identity import DefaultAzureCredential from azure.ai.projects import AIProjectClient from azure.ai.projects.models import CodeInterpreterTool, FilePurpose from pathlib import Path We will then load the variables from the .env file, and create the AIProjectClient and connect to our Azure AI Foundry Project.\n# load environment variables from local .env file load_dotenv() PROJECT_CONNECTION_STRING = os.getenv(\u0026#34;AZURE_AI_AGENT_PROJECT_CONNECTION_STRING\u0026#34;) MODEL_DEPLOYMENT = os.getenv(\u0026#34;AZURE_AI_AGENT_MODEL_DEPLOYMENT_NAME\u0026#34;) # create ai project client project = AIProjectClient.from_connection_string( conn_str=PROJECT_CONNECTION_STRING, credential=DefaultAzureCredential() ) with project_client: We want our agent (Assistant Coach) to be able to analyze and interpret certain Basketball statistics. In this case, we are going to upload a csv file with statistics about 3-point shots from the NBA seasons 1996 until 2020. To achieve that, we will make use of the CodeInterpreterTool.\n# upload a file and add it to the client file = project_client.agents.upload_file_and_poll( file_path=\u0026#34;nba3p.csv\u0026#34;, purpose=FilePurpose.AGENTS ) print(f\u0026#34;Uploaded file, file ID: {file.id}\u0026#34;) # create a code interpreter tool instance referencing the uploaded file code_interpreter = CodeInterpreterTool(file_ids=[file.id]) Now we can define the agent with instructions and tools.\n# create an agent agent = project_client.agents.create_agent( model=\u0026#34;gpt-4o-mini\u0026#34;, name=\u0026#34;assistant-coach-agent\u0026#34;, instructions=\u0026#34;You are a Basketball assistant coach that knows everything about the game of Basketball. You give advice about Basketball rules, training, statistics and strategies.\u0026#34;, tools=code_interpreter.definitions, tool_resources=code_interpreter.resources, ) print(f\u0026#34;Using agent: {agent.name}\u0026#34;) Next, we create a thread, which is basically the conversation between the user and the agent. In the message itself, we define the ask or task for our agent.\n# create a thread with message thread = project_client.agents.create_thread() print(f\u0026#34;Thread created: {thread.id}\u0026#34;) message = project_client.agents.create_message( thread_id=thread.id, role=\u0026#34;user\u0026#34;, content=\u0026#34;Could you please create a bar chart for 3Pointers made vs attempts during the NBA seasons 1996 until 2020 and save it as a .png file?\u0026#34;, ) Almost done, we now have to specify the run or activation part in which the agent will perform our task. Additionally, we have to fetch the output in reversed chronological order, providing a clear view of the most recent interactions first.\n# ask the agent to perform work on the thread run = project_client.agents.create_and_process_run(thread_id=thread.id, agent_id=agent.id) # fetch and print the conversation history including the last message print(\u0026#34;\\nConversation Log:\\n\u0026#34;) messages = project_client.agents.list_messages(thread_id=thread.id) for message_data in reversed(messages.data): last_message_content = message_data.content[-1] print(f\u0026#34;{message_data.role}: {last_message_content.text.value}\\n\u0026#34;) The agent should generate a graph png that we want to look at. Hence, we have to download the file with the following code snippet.\n# fetch any generated files for file_path_annotation in messages.file_path_annotations: project_client.agents.save_file(file_id=file_path_annotation.file_path.file_id, file_name=Path(file_path_annotation.text).name) print(f\u0026#34;File saved as {Path(file_path_annotation.text).name}\u0026#34;) Lastly, we should clean up and delete the agent and the thread.\n# clean up project_client.agents.delete_agent(agent.id) project_client.agents.delete_thread(thread.id) Running the app # Note: The code snippets above should just give you an overview of the Azure AI Agent Service and are not representing a production-grade application. Error handling, logging, config management and reusability are just some things you need to add to your final application code. Nevertheless, you can find the application code and the used files on my GitHub repo here.\nLet\u0026rsquo;s run our code and see the results.\n(.venv) ➜ basketball-ai-agent python main.py Uploaded file, file ID: assistant-J7sVvG3wVZXRkXHwSxELJY Using agent: assistant-coach-agent Thread created: thread_zqMD7857xH4zsKwN6olVvkmx Conversation Log: MessageRole.USER: Could you please create a bar chart for 3Pointers made vs attempts during the NBA seasons 1996 until 2020 and save it as a .png file? MessageRole.AGENT: Let\u0026#39;s start by examining the contents of the uploaded file to understand the data it contains. After that, we can create the bar chart for 3-pointers made versus attempts for the NBA seasons from 1996 to 2020. MessageRole.AGENT: The uploaded data contains the following columns: - **NBASeasons**: The season year. - **3PointersMade**: The number of 3-pointers made. - **3PointersAttempts**: The number of 3-pointers attempted. - **3PointersPercentage**: The percentage of successful 3-point shots. - **3PointersPercentageShareInTotalPoints**: The share of 3-pointers in total points. We\u0026#39;ll create a bar chart comparing the number of 3-pointers made versus the number of 3-pointers attempted for the NBA seasons from 1996 to 2020. Let\u0026#39;s proceed with that. MessageRole.AGENT: Here is the bar chart comparing 3-pointers made versus 3-pointers attempted from the NBA seasons 1996 to 2020. You can download it using the link below: [Download the chart](sandbox:/mnt/data/3_pointers_made_vs_attempts_1996_2020.png) File saved as 3_pointers_made_vs_attempts_1996_2020.png During the execution of our app, we can see the agent and thread getting created in Azure AI Foundry portal.\nAfter the run the agent and the thread will be deleted due to our clean-up code. If you want to keep it for troubleshooting proposes, just comment or remove the clean-up section.\nIn the Agents view we can also see the Code Interpreter got added as action tool for our agent, and we can see the file we have uploaded.\nLet\u0026rsquo;s check the results. The graph generated is correct and shows quite a raise in 3 Pointers made and attempts. Thanks Steph Curry.\nSummary # The Azure AI Agent Service allows you to quickly and with less effort build, deploy and manage AI Agents. The comprehensive out of the box toolkit and the possibility to select different models give Developers the flexibility they need to develop agent-based AI apps that can successfully accomplish complex tasks.\nThe Azure AI Agent Service can also easily be used with multi-agent frameworks such as Semantic Kernel. Stay tuned for another blog post about multi-agent apps.\nResources # AI agents — what they are, and how they’ll change the way we work Introducing Azure AI Agent Service What is Azure AI Agent Service Azure AI Agent standard deployment via bicep template from Azure AI Samples repository Azure AI Agent Code Interpreter What is Azure AI Foundry Azure AI Foundry SDK Azure AI Foundry Hubs and Projects ","date":"2 April 2025","externalUrl":null,"permalink":"/posts/agent/","section":"Posts","summary":"Intro # AI Agents are in talks, and we are going to take a look at the Azure AI Agent Service.","title":"Azure AI Agent Service: Build and Run AI Agents on Azure","type":"posts"},{"content":" Disclaimer # 🚨Please be aware that all content posted on this website represents only my personal opinion and is based on my experience – it does not represent Microsoft\u0026rsquo;s positions, strategies, or opinions.🚨\nPrivacy \u0026amp; analytics # I want to keep this simple and respectful of your privacy. This blog uses Umami, a privacy-friendly, cookieless analytics tool that only records aggregated, anonymous metrics such as page views and referrers. No tracking cookies are set and no personal data is collected about you.\n","date":"30 January 2025","externalUrl":null,"permalink":"/disclaimer/","section":"BeyondElastic","summary":"Disclaimer # 🚨Please be aware that all content posted on this website represents only my personal opinion and is based on my experience – it does not represent Microsoft\u0026rsquo;s positions, strategies, or opinions.","title":"Disclaimer","type":"page"},{"content":"","date":"30 January 2025","externalUrl":null,"permalink":"/tags/intro/","section":"Tags","summary":"","title":"Intro","type":"tags"},{"content":"Hi everyone,\nWelcome to my blog! 🎉 Here, I dive into the exciting world of modern application development, platform challenges, and cloud solutions. As a Technical Specialist, I\u0026rsquo;m thrilled to explore the latest trends and technologies that are shaping our industry.\nIn this space, I\u0026rsquo;ll share insights on how we can tackle today\u0026rsquo;s IT challenges using the powerful capabilities of Microsoft Azure. From cloud computing to AI-driven solutions, DevOps, and everything in between, join me on this journey to discover innovative solutions that drive efficiency and success.\nFor those who have followed my previous blog, beyondelastic.com, thank you for your continued support! I can\u0026rsquo;t wait to bring you even more valuable and fun content here. This is my first post on my new blog, and I\u0026rsquo;m excited to embark on this adventure with you all.\nStay tuned for exciting posts, and don\u0026rsquo;t hesitate to engage with your thoughts and questions!\nBest regards,\nAlex\n","date":"30 January 2025","externalUrl":null,"permalink":"/posts/intro-post/","section":"Posts","summary":"Hi everyone,","title":"Intro Post","type":"posts"},{"content":" $ whoami alexander-ullah $ kubectl get engineer alexander-ullah -o yaml▋ apiVersion: beyondelastic.io/v1 kind: Engineer metadata: name: Alexander Ullah title: Solution Engineer – Cloud \u0026amp; AI Apps company: Microsoft region: EMEA location: Germany labels: experience: \u0026#34;20y+\u0026#34; focus: cloud-native-and-ai spec: about: \u0026gt; GenAI and agentic systems are my focus today. I\u0026#39;m a hands-on, customer-focused technologist who is passionate about AI agents, Kubernetes, cloud-native platforms and modern app architectures. I help customers modernize from legacy to cloud, and share what I learn by blogging and speaking. expertise: - GenAI \u0026amp; multi-agent systems - Kubernetes \u0026amp; containers - Platform engineering - GitOps \u0026amp; software supply chain security - Networking \u0026amp; security - Solution selling \u0026amp; public speaking currentRole: company: Microsoft team: Specialist Team Unit - Cloud \u0026amp; AI period: 2025-01 → present highlights: - GenAI \u0026amp; agentic app strategies for strategic German customers - Agentic DevOps \u0026amp; AI coding assistants to boost developer productivity - Leads MVPs, PoCs, RFPs \u0026amp; hackathons across Azure AI, Azure App \u0026amp; GitHub - Speaker at Agentic AI \u0026amp; DevOps Day previousRole: company: VMware (Tanzu) title: Principal Specialist Solution Engineer \u0026amp; CTO Ambassador period: 2011 → 2024-12 highlights: - App modernization (12-Factor, 5Rs) with the Tanzu portfolio across EMEA - CTO Ambassador bridging R\u0026amp;D and field teams (2023–2024) - Speaker at VMworld 2019–2021 \u0026amp; VMware Explore 2022–2023 - vExpert 2020–2023 \u0026amp; vExpert App Modernization progression: - Senior Lead Solution Engineer, Tanzu (2022–2024) - Lead Solution Engineer, Kubernetes (2019–2022) - Staff Technical Account Manager (2017–2019) - ... certifications: - Azure AI Engineer Associate - Azure Administrator Associate - Azure AI Fundamentals - Certified Kubernetes Administrator (CKA) - Certified Kubernetes Application Developer (CKAD) - Certified Kubernetes Security Specialist (CKS) - VMware vExpert - ... awards: - VMware MVP Specialist SE, CEMEA (Q1 FY23) - VMware \u0026#34;At Our Best\u0026#34; – Elevate Award - VMware \u0026#34;At Our Best\u0026#34; – Achieve Award - ... links: linkedin: https://www.linkedin.com/in/alexander-ullah-a278a0b7/ You can also find me on LinkedIn.\nPlease be aware that all posts and opinions are my own, see Disclaimer\n","date":"30 January 2025","externalUrl":null,"permalink":"/author/","section":"BeyondElastic","summary":"$ whoami alexander-ullah $ kubectl get engineer alexander-ullah -o yaml▋ apiVersion: beyondelastic.","title":"Author","type":"page"},{"content":"","externalUrl":null,"permalink":"/authors/","section":"Authors","summary":"","title":"Authors","type":"authors"},{"content":"","externalUrl":null,"permalink":"/series/","section":"Series","summary":"","title":"Series","type":"series"},{"content":"Hands-on workshops I have built for learning Microsoft AI and Azure technology. Each tile links out to the full workshop you can clone and work through yourself.\n","externalUrl":null,"permalink":"/workshops/","section":"Workshops","summary":"Hands-on workshops I have built for learning Microsoft AI and Azure technology.","title":"Workshops","type":"workshops"}]