How it works

You hold the throttle.

Six straight answers about how you schedule, steer, prove, correct, trace, and cap the AI trains working on your data. Nothing runs, changes, or spends without a control you hold.

Q01 · Dispatch & triggers

Runs fire on demand, on a schedule, or on an event.

Every train starts from one kick-off instruction. It doesn't care who pulls the lever: a person, a clock, or an upstream event. Recurring and event-driven runs are first-class, not bolted on.

ONE ENTRY
A single programmatic kick-off launches any train, the same door a click, a timer, or a webhook uses.
EVENT SPINE
Scheduled jobs already run in production on an event backbone. Cost metering and housekeeping run on timers today.
TAGGED
Every run records how it was triggered, so scheduled and event runs sit in your history beside manual ones.
FROM YOUR SYSTEMS
Connectors listen for activity (a document dropped in your store), so your systems can start a run.
On-demand live · schedules & hooks on the same rail
Dispatch example

A trigger you'd set up

  1. 1New filing lands in your document store
  2. →Connector event fires the review train
  3. →Train runs, scored + logged like any run
  4. →Result posts to your dashboard automatically

Also: “every weekday 6am” or “on close-of-books” attach to the same kick-off.

Q02 · Model choice

No lock-in. Models are configuration, chosen per train, even per railroad car.

The platform is model-agnostic. Leading models from multiple providers sit behind one interface, and which one a train uses is a setting you own, not something wired into code.

MULTI-PROVIDER
Models hosted on Amazon Bedrock plus OpenAI, behind a single routing layer.
PER-CAR
Each train config carries a model block: set a default, then override it for individual railroad cars inside the train.
CONFIG, NOT CODE
Switching a model is a config change and a version bump: no code change, no redeploy.
ALL THE WAY DOWN
Even the model that grades quality (the eval judge) is configurable.
Available now
Mixed motive power

One train, two model tiers

  1. ◆Drafting car → a premium model
  2. ◆Extraction car → a cheaper, faster model
  3. ◆Judge → whichever model you trust to score

Spend the premium budget only where the task needs it.

Q03 · Safe model swaps

You never swap blind. Every change is proven against your evals before it ships.

Moving from an expensive model to a cheaper one is a measured decision, not a gamble. Each train has its own test suite of real, ground-truth questions, and a change can't reach production until it passes.

4-LAYER SCORE
Every run is scored on structure, behavior, correctness, and quality as a pass/fail card, not a gut feel.
BEFORE / AFTER
Point the train at the cheaper model in a draft, run its suite, and see the pass-rate delta versus today, question by question.
THE GATE
Promotion is gated: any failing must-pass case (or a runner that produced nothing) blocks the change from dev → staging → production.
ON THE RECORD
Results are an immutable ledger, each pinned to the exact config version, so “did it regress?” is a fact.
Available now
Test-track run

Swapping to a cheaper model

  1. 1Draft: premium → cheaper model
  2. 2Run the train's eval suite
  3. ✓Correctness holds 96% → 96% → gate opens
  4. ✗Correctness drops to 78% → gate stays shut

The cheaper model only ships if the numbers say it's safe.

Q04 · Feedback & re-teaching

Correct a railroad car in plain language, and make the fix stick.

When a run goes wrong, you don't file a ticket and wait. You point the Conductor at the run, say what was wrong, and it proposes a concrete fix you approve. Then you lock the correction in as a permanent test.

PROPOSE
The Conductor suggests a real edit, shown as a diff: a worked example, a guidance line, tone, domain terms, even the model.
YOU APPROVE
Nothing changes until you say so; publishing is a separate, harder confirm. Every edit is attributed to you and logged.
MAKE IT STICK
Turn “the right answer is X” into a permanent eval case, so the mistake can't quietly return.
SEE IT FIRST
Dry-run the fix and read the before/after scorecard before you publish.
Live in production · every edit human-approved
Re-teaching a train

From bad run to fixed train

  1. “Run #1742 gave the wrong count and rambled. Fix it.”
  2. →Conductor: add an example pinning the right count + a “be concise” line
  3. →You approve the diff · dry-run · scorecard improves
  4. ✓Publish, logged as your change

The assistant never edits more than you allow, and never publishes on its own.

Q05 · Audit & rollback

Every change is versioned, logged, and reversible, with a small blast radius by design.

“Revert” isn't one magic button; it's three guarantees working together: you can see exactly what happened, roll configuration back a version, and trust that railroad cars were never allowed to do unbounded damage in the first place.

FULL RECORD
Every run stores its inputs, outputs, and the exact template + config version behind it. Every approved change writes an append-only audit row naming a person.
ROLL BACK
Train configurations are version-stored: revert to any prior version in one step, and past runs stay reproducible against the version that made them.
SMALL BLAST RADIUS
Railroad cars can't exceed you: outward writes are approved first, connectors enforce allow-lists that fail closed, and avoid hard deletes.
TRACE & UNDO
Because every change is scoped, attributed and logged, a human can see precisely what changed, and undo it.
Versioning · audit · safe writes · available now
Rollback path

Undoing a change

  1. ?A config edit made results worse
  2. →Open the audit trail: who, what, which version
  3. →Restore the prior version, one step
  4. ✓New runs use the restored config immediately

Honest caveat: an action already taken in the outside world (an email sent) is done, which is exactly why outward actions are approved before they happen.

Q06 · Cost & limits

Every token and tool call is metered and attributed to the run, the train, and the person.

You never get a mystery bill. Cost is measured at the finest grain and rolled up the way you actually think about it, so you can see where the money goes and act on it.

FINE-GRAINED
Model usage and tool/connector usage are recorded per run, then rolled up per train, per workload, and per user who triggered it.
TRUE COST
Compute and data-transfer cost are attributed per customer too, so you see real cost, not just token counts.
BUDGETS
Each account carries a monthly budget with 50 / 75 / 90 / 100% alerts; at the ceiling, new runs are held.
ADJUST
Because spend is attributed by train, model and user, you act on it: move a train to a cheaper model (proven safe above), or tighten scope.
Metering & budgets live · automated hard caps hardening
The meter

What you can answer

  1. ▸“What did this train cost this month?”
  2. ▸“Which user or workload drove spend?”
  3. ▸“Are we near the monthly budget?”
  4. ▸“Where can we drop to a cheaper model safely?”

Attribution goes down to the individual person who kicked off a run.

Six questions, one answer: you stay the operator.

A trigger, a config, a gate, an approval, a version, a meter: every one of them is a lever in your hand. Tell us what your team keeps redoing, and we'll show you a train doing it on your data, with these controls wired in from day one.