Production AI & Agentic Systems

Software Engineer @ Microsoft Copilot · 2026–Present

The Problem

A model call in a notebook is a demo. A model call that millions of people depend on every day is a distributed systems problem. Copilot runs on frontier models from OpenAI, Anthropic, and Microsoft AI, wraps them in agents and tools, and grounds them in real data — all under strict latency and reliability budgets. The challenge is making that stack behave predictably at scale: agents that fail gracefully, retrieval that stays relevant, evaluation that catches regressions before users do, and a backend that holds its latency as traffic grows.

What I Do

I work on the production AI systems behind Copilot — the layer between frontier models and the people using them. The work spans agent orchestration, tool integration, retrieval and grounding, evaluation, and the distributed backend that ties it together.

  • Build model-agnostic AI systems that route across frontier models from OpenAI, Anthropic, and Microsoft AI, so the product isn't locked to any single provider
  • Design agentic workflows with tool use — including MCP-based tools and skills — that let models take reliable, observable actions instead of just generating text
  • Improve retrieval and grounding so responses are anchored in the right context, and build evaluation to measure quality continuously rather than by spot-checking
  • Own distributed backend services handling millions of requests per day, engineered for reliability and observability under real production load
  • Harden failure handling and fallbacks so degraded models, slow tools, or bad context don't turn into user-visible outages

The Hard Parts

Production AI breaks the assumptions most backends are built on. Calls are slow, non-deterministic, and cost real money per request. A response can be structurally valid and semantically wrong. Latency and cost become functions of user input, not just traffic. The engineering that matters is everything around the model: budgets and timeouts, graceful degradation, observability into what the model and agents are actually doing, and evaluation that tells you whether a change made things better or worse before it ships.

Impact

  • Reduced hard failures by 88% through better failure handling, fallbacks, and reliability engineering
  • Supported 26% traffic growth without regressing reliability
  • Held P95 latency near one second on AI-backed request paths at scale
  • Helped make Copilot's AI systems model-agnostic, resilient, and observable in production