Adobe Firefly supports creative AI features across Adobe apps by running massive GPU-based training jobs on Amazon EKS. To handle the explosive growth in monitoring data from thousands of GPUs, Adobe migrated critical observability metrics from a self-managed Prometheus setup to Amazon's fully managed Prometheus service, achieving 28 times faster query performance and higher infrastructure resilience.
- 28x faster metric queries for GPU training jobs
- Managed Prometheus reduces operational maintenance
- Incremental migration preserves existing monitoring workflows
Infrastructure signal
Adobe Firefly runs GPU-intensive model training on Amazon Elastic Kubernetes Service (EKS), scaling up to thousands of compute nodes and over 16,000 GPUs. The large cardinality and volume of telemetry data—over 1 billion data points per query window—challenge traditional self-hosted Prometheus systems. Adobe’s self-managed infrastructure struggled to meet the query performance and availability requirements as the Firefly service adoption grew.
To address these issues, Adobe adopted Amazon Managed Service for Prometheus, a fully managed, scalable observability backend designed to handle extremely large metric volumes with high availability. The managed service allows horizontal scalability and eliminates the operational burden of maintaining Prometheus at scale, while integrated managed scrapers ensure data ingestion from EKS clusters remains seamless.
Developer impact
With the migration to Amazon Managed Prometheus, developers and infrastructure engineers gained significantly faster access to performance and health metrics for GPU training workloads. Query speeds increased 28-fold for critical GPU utilization metrics, enabling more timely troubleshooting and optimization. Self-service access was simplified, empowering teams to independently monitor and analyze training jobs without bottlenecks.
By focusing on a curated set of critical metrics, Adobe enabled developers to better understand complex interactions among GPU compute, memory, and network resources, rather than relying on aggregate CPU-centric statistics. This fine-grained observability improves diagnosis of bottlenecks and supports efficient scaling of model training pipelines.
What teams should watch
Teams managing large-scale, high-cardinality GPU infrastructure should consider managed observability services that can horizontally scale and offload operational complexity. Incremental migration strategies, such as running managed scrapers alongside existing Prometheus setups, offer risk mitigation and continuity. Monitoring services like Amazon Managed Prometheus and Managed Grafana do incur additional costs based on ingestion and query volume, so teams must evaluate pricing aligned to their workload patterns.
Platform teams should also prioritize iterative feedback from infrastructure users to identify a focused set of critical metrics that balance comprehensive insight with performance. Such collaborative curation avoids excessive metric noise and improves usability of observability dashboards, supporting faster root cause analysis and deployment confidence in evolving cloud AI infrastructures.