The console, exactly as your team would see it.
Four surfaces answer the four questions that matter: is anything broken, what went wrong, is quality holding, and what does it cost. The trains below are the work where being wrong is expensive: auditing your own code, vetting what customers send you, catching schedule risk before it lands, and grading the assistant you have already put in front of customers.
The departure board
Is anything broken right now?
Every train on one board, attention first, so a fault can never sit below the fold. The red eval alert means a quality check is failing even though runs are succeeding: that's the 'silent regression' the platform exists to catch.
5 trains on the board
Aug 1 – Aug 18, 2026
- Faultironwood-agent-reviewagent_qaagent-answer-review@v0.815d28 runs2 failed15d287%3m
- Runningironwood-doc-intakedoc_intakedoc-intake-review@v1.615d31 runs15d310%6m
- On timeironwood-security-auditevals 1security_auditsecurity-audit@v2.115d24 runs15d240%11m
- On timeironwood-schedule-riskschedule_riskschedule-risk@v1.315d20 runs15d200%2m
- No serviceironwood-commissioningcommissioningcommissioning-pack@v0.437d—0 runs37d————
Run diagnosis
What went wrong overnight?
A failed run opens on the cause, not a log dump: the step timeline below plays back exactly which railroad car called which tool, and where it died. One click re-runs it once the upstream recovers.
DocsGatewayError: manual lookup failed for 14 of 96 graded answers; upstream /v3/articles returned 502 three times (circuit opened after retry budget)Error (full text below)
DocsGatewayError: manual lookup failed for 14 of 96 graded answers; upstream /v3/articles returned 502 three times (circuit opened after retry budget)
- Phase · load-config+0.0s
- Transcript Reader · started+2.0s
- Transcript Reader · fetch_support_transcripts3.8s+4.0s
- Transcript Reader · execute_select2.1s+9.0s
- Manual Grader · started+14.0s
- Manual Grader · docs_article_lookup9.4s+16.0s
- Manual Grader · docs_article_lookup12.6sHTTP 502 from /v3/articles (attempt 3 of 3)+29.0s
- ErrorDocsGatewayError: manual lookup failed for 14 of 96 graded answers (in docs_article_lookup)+43.0s
No silent regressions
Is quality holding, version over version?
Every run is scored against a test suite built for this business, and history is kept per check, so a regression shows up as a red column, not a support ticket. The amber flag marks a flickering check under review.
| Pass rate | 100 | 100 | 100 | 92 | 100 | 100 | 100 | 100 | 92 | 100 | 100 | 100 | 92 | 85 | 92 | 85 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| L1No secret value in the report | ||||||||||||||||
| L1Report matches declared JSON shape | ||||||||||||||||
| L1Response is non-empty | ||||||||||||||||
| L2Dependency manifest parsed | ||||||||||||||||
| L2Every changed file was read | ||||||||||||||||
| L2No write tool called, ever | ||||||||||||||||
| L3Every finding cites file and line | ||||||||||||||||
| L3Every finding is reproducible | ||||||||||||||||
| L3Flags hardcoded credentials | ||||||||||||||||
| L3No duplicate findings | ||||||||||||||||
| L3Severity matches the CVSS band | ||||||||||||||||
| L4Escalates exposure to a human | ||||||||||||||||
| L4Summary reads plainlytracked |
pass fail skipped not evaluated· each column is one evaluated run, oldest → newest
Costed, not estimated
What does it cost, per run, in dollars?
Every model call and tool call is metered and attributed. Not credits, not tokens: dollars, per run, per train, with the pricing provenance to back it.
| Train | Share | Cost |
|---|---|---|
| ironwood-security-audit | 58% | $48.20 |
| ironwood-doc-intake | 23% | $19.30 |
| ironwood-agent-review | 13% | $10.60 |
| ironwood-schedule-risk | 5% | $4.47 |
Priced via sample-2026-08 · authoritative
| Run | Train | Status | Who | Duration | Cost | When |
|---|---|---|---|---|---|---|
| 0199…0041 | ironwood-doc-intake | running | — | — | $0.00 | Aug 18, 1:58 PM |
| 0199…0040 | ironwood-security-audit | succeeded | ci | 11m 48s | $2.14 | Aug 18, 1:02 PM |
| 0199…0039 | ironwood-schedule-risk | succeeded | — | 2m 41s | $0.22 | Aug 17, 11:36 PM |
| 0199…0038 | ironwood-agent-review | failed | sre | 47s | $0.09 | Aug 17, 6:48 PM |
| 0199…0037 | ironwood-security-audit | succeeded | ci | 10m 21s | $1.98 | Aug 17, 1:02 PM |
| 0199…0036 | ironwood-doc-intake | succeeded | — | 6m 5s | $0.63 | Aug 17, 6:48 AM |
| 0199…0035 | ironwood-schedule-risk | succeeded | — | 2m 33s | $0.21 | Aug 16, 11:36 PM |
| 0199…0034 | ironwood-agent-review | succeeded | — | 4m 2s | $0.40 | Aug 16, 1:02 PM |
See it run on your own work.
Tell us what your team keeps redoing, and we’ll show you a train doing it on your data, scored against checks built for your business.
Talk to us →