Oracli · Predictive infrastructure intelligence

Know what your
cloud does next.

Oracli forecasts cost, capacity and reliability risk across your cloud and GPU fleet — finds what is driving it, and drafts the change that prevents it.

Forecast horizon · 90 days AWS · GCP · Azure · CoreWeave Read-only by default
Oracli · control plane run 4127 · live
VP

Our GPU bill is up 38% this month but training throughput hasn’t moved. Find what’s driving it and tell me what to change before the next billing cycle closes.

@spend-trace

Reconciles billing exports, commitment coverage and credits across every account and region.

41 days parsed · $612K delta isolated to 3 SKUs
@fleet-probe

Samples DCGM, Prometheus and scheduler telemetry to see what the hardware was actually doing.

h100 sm_occupancy p50 31% · ranking-model pool
@cause-graph

Correlates the shift against deploys, config changes and queue depth until one origin survives.

traced to dataloader prefetch change merged aug 04
@forecast

Projects run-rate under each candidate fix and ranks them by saving, blast radius and rollout time.

$1.94M/qtr → $1.53M/qtr at 95% CI

Root cause confirmed. Draft PR opened on infra/dataloader-prefetch with a 6-node canary — projected saving $47K/month at current utilisation.

Ask Oracli about your fleet…
Forecast Q4 Explain spike Capacity plan

The problem

Your dashboards are an expensive way to learn what already happened — the invoice is written weeks before the graph turns red.

How it works

Three steps between raw telemetry and a decision you can defend.

Oracli connects to what you already run — billing exports, metrics, schedulers and your infrastructure code — and stays read-only until you give it a repo to write to.

How deployment works
Connectors8 live
aws · cur v2synced 2m
gcp · billing bqsynced 4m
coreweave · h100synced 2m
prometheus / dcgm30s scrape
k8s scheduler eventsstreaming
terraform stateread-only
Step 01 · Connect

Every source of truth on one timeline

Billing exports, hardware telemetry, scheduler events and infrastructure code are reconciled to a single clock, so a cost line and the deploy that caused it sit next to each other instead of in four tools.

Spend forecast · 90dp95
actual 41d · projected 49d
drift +38.4% vs 30d baseline
attributed toranking-model pool
Step 02 · Model

Forecasts with a cause attached

Each projection is decomposed down to the workload, team and commit that moves it, with confidence intervals backtested against your own history. You get a number and the sentence that explains the number.

Remediationpr open
+ dataloader.prefetch_factor: 4
~ num_workers: 2 → 8
~ node_pool: h100x8 → h100x4
canary6 / 48 nodes · 24h
projected saving$47K / month
blast radius1 node pool
Step 03 · Act

A change, not another notification

Oracli writes the Terraform change, the scheduler policy or the reservation order, opens it as a pull request with a canary plan, and then holds its own forecast against what actually happened after rollout.

What we believe

Infrastructure decisions belong in the week that spends the money, not the review that explains it.

Capabilities

What you can ask a system that has read every layer.

Forecasting, attribution and remediation run on the same reconciled dataset, so the answer you get in chat is the same number that appears in the board deck.

Book a demo
Q4 projection · by workload90d
training · h100$1.21M ± 4%
inference · a10g$418K ± 3%
storage · s3 + fsx$96K ± 2%
egress$61K ± 9%
commitment coverage 62% · recommended 81%
on-demand exposure $340K at current plan

Forecasting and capacity planning

Ninety-day projections for spend, GPU-hours and headroom, broken down by workload and team, with reservation and commitment coverage modelled against them before you sign anything.

Copilot · session 88grounded
you › why did eu-west-1 double on tuesday?
oracli › retry storm on feature-store held 214 pods for 6h — $9.4K. retry budget changed in commit 7f2ab1.
you › can it happen again?
oracli › yes — 3 services share that config. draft a guardrail?

A copilot with your fleet in context

Ask in plain language and get an answer grounded in your own billing, telemetry and commit history — with the query, the window and the source rows it used, so the answer can be checked rather than trusted.

Drift monitor4 open
idle h100-hours · 7d1,284
orphaned volumes212 · $3.1K/mo
untagged spend6.4%
forecast vs actual±2.1%
policy gpu-idle > 45m → alert + drain

Guardrails and drift monitoring

Idle accelerators, orphaned storage, untagged spend and budget drift are caught in hours instead of at month close, with policies that can page a human or drain the pool automatically.

What you get

Four things that change in the first week.

One reconciled timeline

Spend, utilisation, scheduler events and deploys on a single clock — the prerequisite everyone skips.

Attribution to the commit

Every anomaly resolves to a workload, an owner and the change that introduced it, not to an account ID.

Remediation as a pull request

Fixes arrive as reviewable diffs with a canary plan and a projected saving, through your existing approvals.

A forecast you can sign

Backtested intervals, scored against actuals every cycle, so finance and engineering argue from one number.

Founder

Suneet Maharana
Suneet Maharana
Founder · Oracli

Built by someone who watched compute get wasted at both ends of the stack.

Suneet studied at IIT Kanpur, where his research work leaned on high-throughput computational screening — the kind of job that occupies a cluster for weeks and fails quietly when one stage stalls. He went on to work on large-model systems at Flipkart, at a scale where a single inefficient data path stops being a detail and becomes a line item.

Oracli comes out of the gap between those two worlds. Monitoring said the job was running. Billing said what it cost, thirty days later. Nothing in between noticed that the two disagreed, or did anything about it while it still mattered.

Most infrastructure teams already hold the data needed to predict their own next incident. It is just sitting in four systems that never speak to each other.

Trust & data

Read-only until you decide otherwise.

Oracli installs with read-only IAM roles and a scoped metrics endpoint. It never needs write access to production to produce a forecast — write scope is granted per repository, only when you want changes drafted as pull requests. Customer telemetry is isolated per tenant, never pooled, and never used to train shared models.

SOC 2 Type II Single-tenant isolation BYOC / in-VPC deploy RBAC + SSO / SAML Read-only IAM No training on your data Audit log export Data residency
Deployment
oracli-agent → your vpc
egress → control plane only
retention → configurable, 30d default
write scope → per-repo, revocable

Get started

Find out what next quarter costs.