Skip to content
KardashevLabs
Live forward test

ERCOT day-ahead Spread Issuance.
Scored in public, every day.

Before delivery each day, each Forecast Model posts P10/P50/P90 of the next 24 hours of RT minus DA spread at 15 ERCOT hubs and Settlement Zones. A Spread Issuance is written once. After the hours settle, we score them against realized real-time prices. Old Forecast Models keep their own scored history; we do not rewrite or fold them into a newer run.

Companion: ERCOT's own day-ahead Demand forecast accuracy, scored separately. That is operator scoring, not a Kardashev Spread Issuance.

New to energy markets? How to read this page
What's being predicted. Texas electricity is priced twice: a price locked the day before (day-ahead) and the live price during the actual hour (real-time). They never quite match. We predict the gap: will the live price come in above or below the locked one, and by how much, for every hour, at 15 locations.
"Our miss vs market's miss." The locked day-ahead price is itself the market's best guess of real-time. So it's our benchmark: if our average error is smaller than the market's, the model adds real information. Both numbers are averages in dollars per megawatt-hour.
"Promise kept" (coverage). Every prediction is a range, like a circle drawn before throwing a dart: "80% of outcomes will land inside." This tile counts how often reality actually landed inside our ranges. Near 80% = the model knows exactly how uncertain it is. Well below = overconfident; well above = uselessly cautious.
Paper trading. A simulated bet, only placed when the model's entire range clears zero, i.e., even its pessimistic case agrees on direction. Most calm days that never happens and no bet is placed; that discipline is a feature. Results shown are after estimated fees, with no real money at stake.
Model versions. The model is retrained periodically as we add better data. Each version gets its own scored track record below, tagged v1, v2, etc, an older version's numbers are never revised or folded into the new one, so you can see exactly how each version performed on its own.

Live track record

v22026-07-092026-08-30current model
Predictions scored
14,805
one per hour per location
Our miss vs market's miss
10.22 / 10.53
avg error in $/MWh, ours first. Smaller than the market's number = we forecast better than the price the market locked in
Promise kept?
81.1%
every forecast is a range we claim catches the real outcome 80% of the time. This is how often it actually did
Paper trading P&L
$195
130 hours traded · 52.3% winners · after $0.75/MWh fees
Cooldown skips
834
hours we stayed flat after a >$40/MWh swing in the prior 2h, even when the model wanted to trade (added 2026-07-25 after a reversal event cost real P&L)
Day (UTC)Node-hours tradedNet P&LCoverageCooldown skips
2026-08-30098.9%
2026-08-29084.7%51
2026-08-28081.9%23
2026-08-272$4368.7%101
2026-08-26064.1%77
2026-08-25091.9%9
2026-08-24080.6%74
2026-08-23048.9%89
2026-08-22092.8%65
2026-08-211$4383.0%81
2026-08-2020$67286.7%31
2026-08-191$492.8%5
2026-08-181-$3987.2%51
2026-08-17079.7%1
2026-08-16088.3%
2026-08-152$2389.7%
2026-08-14071.9%6
2026-08-13071.9%1
2026-08-1210-$9973.6%1
2026-08-1118-$348.9%
2026-08-10022.2%
2026-08-09096.3%8
2026-08-08086.9%42
2026-08-07076.1%17
2026-08-06078.8%
2026-08-05084.4%
2026-08-04073.9%2
2026-08-03089.1%32
2026-08-02080.3%23
2026-08-01075.0%44
v12026-07-082026-08-30
Predictions scored
15,750
one per hour per location
Our miss vs market's miss
9.38 / 10.24
avg error in $/MWh, ours first. Smaller than the market's number = we forecast better than the price the market locked in
Promise kept?
75.9%
every forecast is a range we claim catches the real outcome 80% of the time. This is how often it actually did
Paper trading P&L
$1,966
221 hours traded · 79.2% winners · after $0.75/MWh fees
Cooldown skips
837
hours we stayed flat after a >$40/MWh swing in the prior 2h, even when the model wanted to trade (added 2026-07-25 after a reversal event cost real P&L)
Day (UTC)Node-hours tradedNet P&LCoverageCooldown skips
2026-08-30099.3%
2026-08-29078.3%51
2026-08-28085.0%23
2026-08-277$13658.3%101
2026-08-26052.2%79
2026-08-251-$083.6%9
2026-08-24071.9%74
2026-08-23040.0%89
2026-08-22091.1%65
2026-08-211$4361.5%81
2026-08-2015$34990.4%31
2026-08-1911$12982.8%6
2026-08-181-$6084.4%51
2026-08-17081.5%1
2026-08-16078.6%
2026-08-152$2083.9%
2026-08-14068.1%6
2026-08-133$1580.6%1
2026-08-122$1280.3%1
2026-08-119$2644.2%
2026-08-10075.6%
2026-08-09088.6%8
2026-08-08079.7%42
2026-08-07067.5%17
2026-08-061$1676.7%
2026-08-05079.4%
2026-08-042$1068.6%2
2026-08-03080.0%32
2026-08-0214$5679.7%23
2026-08-0125$33974.4%44

Nodal basis track

Separate from the hub TFTs. 30 resource settlement points, each the v1 hub issuance plus hour-of-day quantiles of that node's realized basis versus its hub. Not a new Forecast Model architecture. Hubs above stay the public demo.

basis2026-08-232026-08-30
Predictions scored
5,340
one per hour per location
Our miss vs market's miss
28.86 / 28.41
avg error in $/MWh, ours first. Smaller than the market's number = we forecast better than the price the market locked in
Promise kept?
75.1%
every forecast is a range we claim catches the real outcome 80% of the time. This is how often it actually did
Paper trading P&L
$264
33 hours traded · 66.7% winners · after $0.75/MWh fees
Cooldown skips
914
hours we stayed flat after a >$40/MWh swing in the prior 2h, even when the model wanted to trade (added 2026-07-25 after a reversal event cost real P&L)
Day (UTC)Node-hours tradedNet P&LCoverageCooldown skips
2026-08-301$10100.0%7
2026-08-293$3383.3%112
2026-08-281$10684.9%92
2026-08-2711-$10265.1%216
2026-08-261$2060.7%168
2026-08-252$5382.1%87
2026-08-243$5474.2%198
2026-08-2311$9045.2%34

Explore the calls

Pick a node. See already-scored history, or the live, unresolved call for the next 24 hours. Nothing in Live mode has happened yet.

already scored against reality
Pick a node
Window

The shaded area is the range the model claims 80% of outcomes will land in; the thin line inside it is its single best guess. The dots are what actually happened: amber if it landed inside the claimed range, red if the model missed.

Loading HB_HOUSTON

Methodology

The model is a global temporal fusion transformer trained on hourly ERCOT settlement prices from 2019 onward (self-collected from ERCOT MIS archives) plus ERCOT load, day-ahead load forecast, wind and solar generation from EIA-930. The current version (v2) adds ERCOT's own day-ahead wind/solar forecasts, planned outage capacity, ancillary clearing prices and natural gas price as further leak-free inputs. It forecasts the distribution of the next 24 hours' RT−DA spread per node and is retrained quarterly.

The setup is leak-free: at issuance the model sees only information available before delivery, the cleared day-ahead price, published forecasts, and real-time history through the last settled hour. The nodal basis track uses the same rule: its hour-of-day quantiles are fit only on hours already settled.

The paper trading rule is deliberately simple: long the spread when P10 > 0, short when P90 < 0, flat otherwise, 1 MWh per node-hour, haircut $0.75/MWh for fees. In backtest v2 beat v1 on both accuracy (MAE, RMSE) and trading P&L on the same held-out year, with uncertainty bands calibrated using a conformal widening fit on a separate calibration window.

Backtests can flatter; that is what this page is for. Paper fills at hub settlement prices, modeled fees, no market impact. Judge the live numbers.