A leading Asian fashion e-commerce platform leverages Databricks and Lakebase to ingest 1,000+ events per second and deliver highly personalized product recommendations in real time. This architecture balances offline batch pre-computation with real-time API inference to optimize latency, reliability, and developer productivity.
- 1,000+ event/sec ingestion with zero-latency API inference
- Unified feature management ensures training-serving parity
- Dual-path serving architecture balances batch precomputation with dynamic realtime
Infrastructure signal
The architecture employs a unified lakehouse platform on Databricks that ingests shopping behavior events via Lakeflow Connect’s Zerobus Ingest directly into governed Delta tables. This approach removes the need for managing separate message queue infrastructure such as Kafka brokers, reducing overhead and streamlining operational complexity. The event data lands in a bronze layer as raw append-only streams supporting offline feature engineering and ML training.
To achieve low-latency serving, the system uses a dual-path solution: pre-computed batch recommendations serve predictable, high-volume surfaces, while real-time inference endpoints process in-session user signals included directly in the API requests. This design bypasses lakehouse storage for session data during inference, minimizing latency and cost. Unity Catalog provides fine-grained governance across all data layers, protecting sensitive user information while enabling compliance and auditability.
Developer impact
Developers benefit from a consistent feature management framework that ties offline training features to online serving features through the Databricks Feature Store. This eliminates the common ML challenge of training-serving skew and accelerates iteration cycles by using the same definitions and data pipelines for both model development and production inference.
Databricks Workflows orchestrate model training and batch scoring pipelines to refresh behavioral and embedding features at appropriate cadences — daily for user/item embeddings and weekly for product catalogs. This reduces manual operational burden for data scientists and engineers, supporting continuous model improvement without service interruption or manual intervention.
What teams should watch
Data engineering and platform teams should monitor the dual-path serving architecture closely, ensuring that the API-based real-time inference path remains performant and scalable as user signals increase. Performance testing around the direct ingestion bypass of lakehouse storage will be critical to avoid unexpected latency spikes.
Model operations and product teams should track feature pipeline freshness and data quality in the bronze and gold layers managed by Unity Catalog and the Feature Store. Maintaining tight integration between feature refresh schedules and model retraining pipelines is essential to sustaining recommendation relevance and conversion uplift.
Security and compliance teams must leverage Unity Catalog’s lineage and access controls to validate governance over user PII and aggregated feature data. As personalization efforts grow, ensuring data protection compliance while enabling rich feature availability for model training remains a priority.