CumulusBlog
Apr
Mar
From If-Else to Declared Routing
Most production AI systems route between providers with hand-written if-else statements that never get tested. Declared routing rules fix the failover problem before it pages you at 3 AM.
SFT and Online RL for Visual Generation: How We Built CoSprite's Training Pipeline
How Cumulus built a production pipeline for consistent AI-generated game previews using best-of-N sampling, deterministic rendering, pairwise judging, supervised fine-tuning, and online reinforcement learning with GRPO.
Feb
Drop-In OpenAI: Migrating in One Line
How to move an existing OpenAI-SDK application to Cumulus without changing your code — and what you get from the routing, caching, and observability layers once you do.
5 VLMs, 1 GPU: Beating Together AI on Price and Throughput
Cheap GPU inference for AI models: we ran 5 VLMs on one GPU and matched Together AI's throughput at a fraction of the cost. Serverless GPU vs dedicated GPU economics.
Continuous Evaluation: Beyond Judge Consensus
LLM-as-judge with majority vote is the default — and it is wrong often enough to lose trust. A better approach combines synthetic data, deterministic heuristics, judges, and shadow evaluation on live traffic.
Why We Built a Cheaper, Faster GPU Cloud for AI Model Hosting
Cumulus Labs is building the cheapest serverless GPU cloud for AI model hosting. Here's why dedicated GPU instances waste money and how pay-per-second GPU inference changes the economics.




