Faced with surging volumes and expanding product categories, Zepto built a dual-loop evaluation infrastructure on Databricks and MLflow that shifts AI agent development from rapid shipping to rigorous continuous validation, ensuring reliable real-time support across 100,000+ daily tickets.

  • Evaluation-first architecture enables continuous improvement of AI agents
  • MLflow-powered granular tracing ensures full workflow observability
  • Dual development and production loops enhance reliability and control

Infrastructure signal

Zepto’s customer support platform processes over 100,000 AI-driven tickets daily, requiring low latency and high reliability. To manage this, the company built an AI infrastructure centered around Databricks and MLflow, enabling detailed execution tracing and standardized evaluation. Automated tracing captures prompts, responses, tool calls, and decision paths as OpenTelemetry spans, centralized in Delta tables via Unity Catalog.

This infrastructure underpins an evaluation framework that connects development and production cycles through a quality gate. The system's layered observability lets engineers detect failures at any step in the multi-stage agent workflows, not just in final output, improving overall stability and reducing costly error rates during traffic spikes caused by seasonal events or category expansions.

Developer impact

Developers at Zepto shifted from a "just ship" mindset to an evaluation-first process where every AI agent iteration is rigorously measured against shared success criteria defined by all stakeholders. The framework enforces quality gates based on multiple pillars such as customer experience, operational efficiency, and compliance, transforming subjective quality debates into numeric thresholds that must be met prior to deployment.

The introduction of dual feedback loops—one for development and one for production—enables rapid identification and automated correction of failures. This approach ensures agents are not only built confidently but maintained under continuous improvement cycles, minimizing surprise outages and deployment rollback risks while accelerating innovation velocity.

What teams should watch

Teams implementing AI agents for real-time customer service should prioritize building robust evaluation frameworks that integrate observability tools like MLflow and OpenTelemetry to capture full end-to-end execution traces. Defining consensus evaluation pillars that reflect diverse stakeholder goals ensures deployments meet operational, financial, and compliance standards.

Additionally, adopting a dual-loop model linking pre-production evaluation with production monitoring and failure feedback can significantly improve agent reliability at scale. As Zepto’s case demonstrates, this evaluation-first approach reduces costly failure scenarios, enables handling heterogeneous and multilingual customer requests, and supports rapid category expansion without degrading customer support experience.

Source assisted: This briefing began from a discovered source item from Databricks Blog. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings