About
Cedana builds live-migration infrastructure for CPU and GPU workloads, serving AI training and inference, HPC, DevTools, MLOps, and computational-biology users. Its differentiator is transparent, no-code save/migrate/resume functionality that preserves workload state through failures, spot-instance reclamation, and hardware changes while improving utilization and reducing cost.
Market
Cedana competes in AI/GPU compute infrastructure and orchestration, positioning itself as a system-level GPU job-migration layer for AI factories and shared compute operators rather than only a scheduler, cluster manager, or GPU cloud. Its differentiation is full-state, live Save/Migrate/Resume across nodes, clusters, instances, regions, and clouds—including memory, KV cache, CUDA context, and containers—with no code or configuration changes, enabling elastic inference, spot recovery, higher utilization, and workload-level reliability on customer infrastructure.
Cedana primarily serves GPU-intensive infrastructure operators and organizations running shared or distributed AI/HPC workloads, including inference providers, neoclouds and GPU clouds, digital-native and consumer-tech companies, research-computing groups, and pharma/biotech R&D teams. Its likely buyers are platform and infrastructure leaders, GPU-fleet operators, ML-serving teams, and research-computing leaders at growth-stage to enterprise organizations that need higher utilization, lower inference cost, elastic capacity, and resilience to failures or spot-instance reclamation.
At a Glance
Problem
Cedana addresses the high cost and poor utilization of AI and HPC compute, especially GPUs. GPUs are expensive but frequently idle or stranded by static scheduling: Cedana cites average cluster activity below 60% and GPU utilization that rarely exceeds 30%, with fragmentation creating wasted capacity, higher inference costs, manual overhead, and delayed research. The problem is compounded by failures, maintenance, and spot-instance reclamation, which can force long-running jobs to restart from scratch, while spiky inference demand makes keeping fully loaded models warm economically wasteful.
The central use case is a large GPU cluster running both inference and training. When inference demand rises, lower-priority training can be paused or moved rather than forcing the operator to buy peak capacity; when a GPU fails or a spot instance is reclaimed, the workload can resume without losing progress. Cedana reports that useful-work time on customer clusters can rise from roughly 30–40% to more than 80%, and that fully loaded models can be brought online much faster than through a conventional cold start.
Product / Service
Cedana is a compute-infrastructure layer that continuously saves, migrates, and resumes live CPU and GPU workloads. Its checkpoint-and-restore primitive moves running processes or containers between instances, nodes, clusters, or vendors without restarting the job or changing application code. It is designed to work alongside existing Kubernetes, Slurm, Kueue, Ray, Armada, KServe, and NVIDIA Dynamo environments, so customers can add workload mobility without replacing their scheduler or rewriting workflows.
The product is offered both as an open-source package and as a managed service, with infrastructure provisioned and managed through the customer's existing credentials. The benefit is elastic, failure-resilient compute: operators can repack clusters as demand and capacity change, reclaim idle GPUs, perform maintenance without losing progress, and reduce time to serve models. Cedana's published claims include savings of up to 80%, two-to-ten-times faster time to first token, and a 2.28-times faster resume than a native cold start in one 8× H100 benchmark.
Market
Cedana competes in AI infrastructure and cloud-compute orchestration, more specifically in GPU workload management, live migration, utilization optimization, and reliability. Its adjacent competitive set includes incumbent schedulers such as Slurm, Kubernetes, Kueue, and Ray, which Cedana describes as unable to move already-running jobs, as well as GPU-orchestration platforms such as NVIDIA Run:ai and SkyPilot. Cedana's positioning is somewhat complementary to the incumbent schedulers: it extends them with save-migrate-resume capabilities rather than asking customers to discard them.
The public evidence indicates an active, early commercial company rather than a confirmed pre-revenue startup. Y Combinator lists Cedana as a Summer 2023 company and active, and says its current use cases and customers span AI training and inference, HPC, DevTools, ML-ops platforms, and computational biology. Cedana also reports deployments or results involving real customer clusters, Fortune 100 firms, academic centers, AI research labs, and Caltech. No verified revenue figure is provided in the available evidence, so the strongest conclusion is that Cedana has early customer traction but its revenue scale remains undisclosed.
Founders & Leadership
Funding History
Y Combinator
Recent News
Analytics Insight identified Cedana as a cloud-orchestration startup focused on AI infrastructure. Its technology enables CPU and GPU workloads to move across systems without disruption.
Cedana reported that it can restore models in 57–70 seconds, achieving a peak 29.3× speedup. The benchmark compares this performance with native startup times of roughly 50 minutes across fixed resolutions.
Cedana’s YC company description presents its technology as automated GPU-checkpointing infrastructure for improving AI and HPC cluster utilization and reliability. The role listing also references workload migration involving SLURM and NVIDIA Dynamo.
Cedana described its product as saving, migrating, and resuming CPU and GPU workloads across nodes or clusters without manual checkpoint hooks.
Cedana’s documentation highlighted an API for integrating its workload-migration capability into clusters, with applications in high-performance computing, numerical simulation, and AI training or execution.
Cedana argued that AI and HPC schedulers are fundamentally limited because they cannot migrate running jobs. Its CPU and GPU migration capability is presented as a way to overcome that limit and approach substantially higher utilization.
Active Roles
2Business Model
Cedana appears to monetize B2B infrastructure software through enterprise, demo-led sales of its workload migration and orchestration platform. Public materials do not disclose a specific subscription, usage-based, or licensing formula; third-party data estimates approximately $1.4 million in 2024 revenue.