Sample dataThis is the real console UI over a fabricated tenant, “Ironwood Energy Systems,” a grid-storage builder that writes its own fleet software and ships an assistant to the utilities it serves. Every component below is the one customers use. The trains and runs are invented, so the ids here don’t open anything; in your console they do.

The console, exactly as your team would see it.

Four surfaces answer the four questions that matter: is anything broken, what went wrong, is quality holding, and what does it cost. The trains below are the work where being wrong is expensive: auditing your own code, vetting what customers send you, catching schedule risk before it lands, and grading the assistant you have already put in front of customers.

The departure board

Is anything broken right now?

Every train on one board, attention first, so a fault can never sit below the fold. The red eval alert means a quality check is failing even though runs are succeeding: that's the 'silent regression' the platform exists to catch.

Departure board· This month · so far

5 trains on the board

Aug 1 – Aug 18, 2026

Running
1
Attention
1
Eval alerts
1
Runs
103
Failure rate
2%
Spend
$82.57
Status
Train
Median
  • Fault
    ironwood-agent-review
    agent_qa
    15d28 runs2 failed
    3m
  • Running
    ironwood-doc-intake
    doc_intake
    15d31 runs
    6m
  • On time
    ironwood-security-auditevals 1
    security_audit
    15d24 runs
    11m
  • On time
    ironwood-schedule-risk
    schedule_risk
    15d20 runs
    2m
  • No service
    ironwood-commissioning
    commissioning
    37d—0 runs
    —
Aug 1 – Aug 18, 2026 · trains run on demand, so the board shows the last departure, not a timetable

Run diagnosis

What went wrong overnight?

A failed run opens on the cause, not a log dump: the step timeline below plays back exactly which railroad car called which tool, and where it died. One click re-runs it once the upstream recovers.

Run
0199…0038
failedtrain ironwood-agent-reviewAug 17, 6:48 PM47scost $0.09
by schedule:nightly · scheduled
DocsGatewayError: manual lookup failed for 14 of 96 graded answers; upstream /v3/articles returned 502 three times (circuit opened after retry budget)
DocsGatewayError: manual lookup failed for 14 of 96 graded answers; upstream /v3/articles returned 502 three times (circuit opened after retry budget)
Step timeline
8 events
  1. Phase · load-config+0.0s
  2. Transcript Reader · started+2.0s
  3. Transcript Reader · fetch_support_transcripts3.8s+4.0s
  4. Transcript Reader · execute_select2.1s+9.0s
  5. Manual Grader · started+14.0s
  6. Manual Grader · docs_article_lookup9.4s+16.0s
  7. Manual Grader · docs_article_lookup12.6sHTTP 502 from /v3/articles (attempt 3 of 3)+29.0s
  8. ErrorDocsGatewayError: manual lookup failed for 14 of 96 graded answers (in docs_article_lookup)+43.0s

No silent regressions

Is quality holding, version over version?

Every run is scored against a test suite built for this business, and history is kept per check, so a regression shows up as a red column, not a support ticket. The amber flag marks a flickering check under review.

Regression history
suite v9 · last 16 evaluated runs
85%latest pass rate1 regressed·0 improvedvs previous evaluated run
Pass rate100100100921001001001009210010010092859285
L1No secret value in the report
L1Report matches declared JSON shape
L1Response is non-empty
L2Dependency manifest parsed
L2Every changed file was read
L2No write tool called, ever
L3Every finding cites file and line
L3Every finding is reproducible
L3Flags hardcoded credentials
L3No duplicate findings
L3Severity matches the CVSS band
L4Escalates exposure to a human
L4Summary reads plainlytracked

pass fail skipped not evaluated· each column is one evaluated run, oldest → newest

Eval scorecard
0199…0115 fail
L1 Structural3/3100%
L2 Behavioral3/3100%
L3 Outcome3/560%
L4 Quality2/2100%

Costed, not estimated

What does it cost, per run, in dollars?

Every model call and tool call is metered and attributed. Not credits, not tokens: dollars, per run, per train, with the pricing provenance to back it.

Cost
This month
Total spend
$82.57
≈ $0.80 / run · 103 runs
LLM
$75.20
MCP
$4.02
By workload
train-run$71.05eval$4.15connector$4.02infrastructure$3.35
TrainShareCost
ironwood-security-audit58%$48.20
ironwood-doc-intake23%$19.30
ironwood-agent-review13%$10.60
ironwood-schedule-risk5%$4.47

Priced via sample-2026-08 · authoritative

Spend by daysample month
Runs
recent runs · sample
RunTrainStatusWhoDurationCostWhen
0199…0041ironwood-doc-intakerunning——$0.00Aug 18, 1:58 PM
0199…0040ironwood-security-auditsucceededci11m 48s$2.14Aug 18, 1:02 PM
0199…0039ironwood-schedule-risksucceeded—2m 41s$0.22Aug 17, 11:36 PM
0199…0038ironwood-agent-reviewfailedsre47s$0.09Aug 17, 6:48 PM
0199…0037ironwood-security-auditsucceededci10m 21s$1.98Aug 17, 1:02 PM
0199…0036ironwood-doc-intakesucceeded—6m 5s$0.63Aug 17, 6:48 AM
0199…0035ironwood-schedule-risksucceeded—2m 33s$0.21Aug 16, 11:36 PM
0199…0034ironwood-agent-reviewsucceeded—4m 2s$0.40Aug 16, 1:02 PM

See it run on your own work.

Tell us what your team keeps redoing, and we’ll show you a train doing it on your data, scored against checks built for your business.

Talk to us →