Every source of truth on one timeline
Billing exports, hardware telemetry, scheduler events and infrastructure code are reconciled to a single clock, so a cost line and the deploy that caused it sit next to each other instead of in four tools.
Oracli · Predictive infrastructure intelligence
Oracli forecasts cost, capacity and reliability risk across your cloud and GPU fleet — finds what is driving it, and drafts the change that prevents it.
Our GPU bill is up 38% this month but training throughput hasn’t moved. Find what’s driving it and tell me what to change before the next billing cycle closes.
Reconciles billing exports, commitment coverage and credits across every account and region.
Samples DCGM, Prometheus and scheduler telemetry to see what the hardware was actually doing.
Correlates the shift against deploys, config changes and queue depth until one origin survives.
Projects run-rate under each candidate fix and ranks them by saving, blast radius and rollout time.
Root cause confirmed. Draft PR opened on infra/dataloader-prefetch with a 6-node canary — projected saving $47K/month at current utilisation.
The problem
How it works
Oracli connects to what you already run — billing exports, metrics, schedulers and your infrastructure code — and stays read-only until you give it a repo to write to.
How deployment worksBilling exports, hardware telemetry, scheduler events and infrastructure code are reconciled to a single clock, so a cost line and the deploy that caused it sit next to each other instead of in four tools.
Each projection is decomposed down to the workload, team and commit that moves it, with confidence intervals backtested against your own history. You get a number and the sentence that explains the number.
Oracli writes the Terraform change, the scheduler policy or the reservation order, opens it as a pull request with a canary plan, and then holds its own forecast against what actually happened after rollout.
What we believe
Capabilities
Forecasting, attribution and remediation run on the same reconciled dataset, so the answer you get in chat is the same number that appears in the board deck.
Book a demoNinety-day projections for spend, GPU-hours and headroom, broken down by workload and team, with reservation and commitment coverage modelled against them before you sign anything.
Ask in plain language and get an answer grounded in your own billing, telemetry and commit history — with the query, the window and the source rows it used, so the answer can be checked rather than trusted.
Idle accelerators, orphaned storage, untagged spend and budget drift are caught in hours instead of at month close, with policies that can page a human or drain the pool automatically.
What you get
Spend, utilisation, scheduler events and deploys on a single clock — the prerequisite everyone skips.
Every anomaly resolves to a workload, an owner and the change that introduced it, not to an account ID.
Fixes arrive as reviewable diffs with a canary plan and a projected saving, through your existing approvals.
Backtested intervals, scored against actuals every cycle, so finance and engineering argue from one number.
Founder
Suneet studied at IIT Kanpur, where his research work leaned on high-throughput computational screening — the kind of job that occupies a cluster for weeks and fails quietly when one stage stalls. He went on to work on large-model systems at Flipkart, at a scale where a single inefficient data path stops being a detail and becomes a line item.
Oracli comes out of the gap between those two worlds. Monitoring said the job was running. Billing said what it cost, thirty days later. Nothing in between noticed that the two disagreed, or did anything about it while it still mattered.
Most infrastructure teams already hold the data needed to predict their own next incident. It is just sitting in four systems that never speak to each other.
Trust & data
Oracli installs with read-only IAM roles and a scoped metrics endpoint. It never needs write access to production to produce a forecast — write scope is granted per repository, only when you want changes drafted as pull requests. Customer telemetry is isolated per tenant, never pooled, and never used to train shared models.