Skip to content

How we forecast your spend

The “Projected this month / next month” figures on the Costs page (and the budget_forecast_breach / forecast_increase alerts) all come from one forecast, computed only from your own billing history. There is no second model anywhere in the product, and nothing in the number is a canned constant.

For the month in progress, the projection is the highest of three honest readings: we surface risk early and never project below what you have already spent:

  1. Month-to-date: what the bill already says.
  2. Run rate: month-to-date plus the remaining days at the median day of the month so far. Some fees are booked in a lump on day 1 (tax, a support plan, a security bundle). They are in month-to-date once, as booked, and the median ignores them, so a large first day is not multiplied by the length of the month. The “so far” ends at the last day your cost feed reaches, not today’s date. With fewer than three observed days the median is noise, and the run rate is month-to-date divided by the fraction of the month the data covers. A day inside the cut with no charge counts as a zero day, so a fee booked once on day 1 projects as booked. A series that bills on only a few days a month (a weekly job) also looks at its own previous month: the run rate is the larger of the median-day estimate and month to date plus what that series billed after the same day last month.
  3. Model: a Holt-Winters forecast (level + trend + yearly seasonality, smoothing parameters grid-searched per org) fitted on your closed months only. The partial current month never contaminates the training data.

For example, one production tenant billed $41.6k in the first seven days of October, $17.8k of it on day 1. A straight line projects $184k. The run rate above projects $141k.

Next month’s figure is the same model’s two-step-ahead point, published only when its band is at most 30% of the figure. A wider band says “anything”, so the point is withheld and the response says why. The month in progress keeps its projection.

The model trains on months that describe the same estate. When a cloud’s cost history starts mid-window (for example an AWS connector whose history begins in June, so the monthly total jumps 9x while the other cloud runs flat), the jump is a coverage change, not growth, and a model fitted across it would project the new level into everything after. The model uses only the months from the latest start among the clouds in view. With fewer than six such months it drops out and the run rate is the projection. Under a single-cloud filter only that cloud’s start counts.

The band is not a fixed percentage. We replay the model over your own history, a rolling backtest: for every month we have, refit on only the months before it and score the prediction against what actually happened. The ± band is the 80th percentile of those absolute errors, so roughly 8 out of 10 of the model’s own historical misses fell inside it. The band’s lower edge never dips below what you have already spent this month.

The same backtest produces the accuracy figure shown next to the projection: the mean absolute percentage error (MAPE) one month ahead, with the number of backtest samples it was measured on. It answers “how wrong has this forecast typically been for us”, not a benchmark, not a promise.

A payment billed once, such as an annual AWS Marketplace contract, is not a month of normal spend. It stays in every total, so the bill reconciles, but the forecast is fitted without it: one lump in the history would otherwise be carried into every projected month. When the open month holds such a payment, it is added back onto this month’s projection, so the month-end figure still matches the bill. The Next month tile names what was left out (“Excludes a one-time payment of $X in ()”), and the API returns it as projection.excludedOneTime. Service rows in the evolution grid follow the same rule. The “same day last month” reading below has no charge type to go by, so a payment leaves it only as a lump (a day above twice the month’s median day).

A service row whose first billed month opened after day 1 (a connector that started on the 8th) leaves that partial month out of its model, because a model fitted across it reads a ramp.

  • Fewer than 3 closed months: no forecast at all. A trend line through one or two points is a guess, so the surface hides and the API answers with an explicit “not enough history” reason.
  • Fewer than 6 closed months: point forecast only, no ± band and no accuracy figure. A band fitted on one or two residuals is noise dressed up as confidence.
  • Stale history: if the newest closed month on record is older than last month (for example a cloud sync that stopped weeks ago), we refuse with an explicit “history is stale” reason instead of relabelling an old model point as “this month”.

Short histories make wide, noisy bands; we show them as they are rather than tightening them cosmetically. Every number traces to your own cost rows.

The month in progress is never compared as a finished one

Section titled “The month in progress is never compared as a finished one”

A month that is still accruing is not comparable to a closed month, so nothing on the Costs page pretends otherwise. Three treatments, depending on what you are looking at:

  • Any comparison: a month-over-month move, a delta against your baseline, a per-cell change: is measured on the projected month-end and labelled projected. Where there is no defendable projection you get month-to-date with no comparison at all, never a delta on a half-finished month.
  • Any line or bar series runs solid through the last closed month, dashes from there to the projected point, and marks month-to-date hollow.
  • Any windowed table (Kubernetes namespaces, commitment coverage) names its window on screen (last month + this month to date rather than “2 months”) and shows no delta.

The still-accruing column is tagged mtd wherever it appears, and the baseline pickers offer closed months only: a baseline is a yardstick, so it has to be a finished month.

The projection is not only an organization-level number. The same forecast (the same model, the same gates, the same disclosed band) is attached to every row that already has a month-by-month history:

  • each subscription, account or project on the trend matrix;
  • each service or tag value in the evolution grid;
  • each value in the cost-by-tag matrix, over the range you have selected;
  • each resource, in its drawer;
  • each Kubernetes namespace, in the currency its cloud billed.

A service row has its own day-by-day figures, so its run rate is built the same way: month to date plus the remaining days at its median day. A fee booked once on day 1 (a support plan, a security bundle) projects as booked. A row with no spend this month has no projection (“No spend this month yet.”): a service that stopped billing is not projected at last month’s rate. Rows without day-by-day figures (a tag value, a subscription) get one more rule: in the first 10 days of the month, a row that already holds at least 90% of its last closed month projects flat, at what is booked, with no band.

The gates are that row’s own. A row with fewer than three closed months of its own history, or whose history has gone stale, says so in place of a number, and a row that only started billing recently is not credited with the empty months before it existed. A cell that cannot be forecast shows a dash and the reason, never a zero and never a comparison.

Fitting a model is not free, so each view forecasts its 40 largest rows by spend and tells you plainly when a row falls below that line. Those are the rows at the top of the table you are already reading.

Under the projected figure, the Costs headline can also tell you what the same stretch of last month had cost by now. Both sides are cut at the same day of the month, and that day comes from how far your current month’s data actually reaches, not from today’s date.

That distinction is the whole point. Cloud cost data lands a day or two late, so on the 15th your current month typically holds data through the 13th while last month’s 15th is complete. Comparing those two would show a fall every single day of every month. The headline prints the day it used, so you can see the comparison is like-for-like.

A day only counts once it has settled: its own total reaches at least half of last month’s typical day, and the month’s spend through it reaches at least half of last month’s spend through the same day. Cloud feeds backfill out of order, so a few dollars on the 12th, or one normal-looking day amid ten empty ones, is not a landed month and would read as a collapse that is not real.

Both sides are usage, not the bill. A fee booked in a lump differs between months for reasons that are not a change in use, so each side drops the lump excess of its days: for every day above twice that month’s median day, what it holds over the median. The page shows the amounts removed. On the tenant above the raw reading was $41.6k against $56.9k (-27%, almost all of it day 1), and the usage reading is $28.0k against $28.7k (-2.5%).

We leave the comparison out rather than approximate it: while no day of the current month has settled, when last month has none to compare against, and when the page is filtered in a way the daily figures cannot reproduce: a region filter, a cohort lock, or the planned lens. It is an organization-level reading only.