Product

From a blank canvas to a scheduled production pipeline.

Build on a canvas, see the columns before anything runs, test on fifty rows, then publish to staging and promote to production. Here is each step, and the engine underneath all of them.

01 · Build

A canvas, a ribbon of cards, and a form for every one of them.

Drag a node out of the ribbon, drop it on the canvas and wire it up. Each node opens a configuration panel written specifically for it, with the real columns from upstream already in the pickers. You will not find a generic key/value editor anywhere.

Search every node

One search box covers every node, matching name, category and description. It ignores which tab you are on, so you never have to know where a node lives.

Help before you commit

Every node carries its own documentation: what it does, when to reach for it, worked examples and the gotchas. You can read it from the ribbon before you drop the node and from the config panel afterwards.

Annotate the graph

Pin notes to the canvas, and group nodes into a Container you can switch off in one click. The data then passes straight through the section as if it were not there.

Metadata · Join NO RUN NEEDED
order_idInt64
customer_nameUtf8
placed_atTimestamp
amountDecimal128
regionUtf8
is_priorityBoolean
On BigQuery it goes further. A dry run returns the schema and the bytes your query will scan, so the cost of a query is on screen before you have spent a cent running it.
02 · See the schema

Stop running the pipeline just to find out what the columns are.

Trinium walks your graph and propagates types the same way the executor propagates data, without moving a single row. Parquet footers, CSV headers, database catalogs and warehouse metadata answer the question for most sources before anything runs.

Column pickers that are never empty

Every field that takes a column name offers the real ones, typed, the moment the node is wired.

Expressions checked as you type

A formula is validated against the real inbound schema. A misspelled column or a type mismatch shows up in the builder instead of as a failed run at 3am.

It degrades, it never blocks

If one node cannot be known without running, only that branch goes quiet. The rest of the canvas stays lit.

03 · Express

One expression language. It's SQL, and it's the same everywhere.

Filter conditions, Formula columns, multi-field rewrites, window functions and validation rules all use the same syntax. There is no proprietary dialect to learn and no second language hiding in a different node.

Bracket columns still work

Write [unit price] if that is your muscle memory, or "unit price", or plain price. All three resolve to the same column.

A curated reference, not a firehose

Logical, text, number, date and conversion functions, each one documented and tested. Browse them inside the expression field or on a help page in the app.

Cross-row work is just a window function

Running totals, lag and lead, moving averages, ranking and tiling are all standard SQL, so there is no special node vocabulary to memorise.

Formula · new column "band"
-- graded, with a fallback, in one expression
CASE
  WHEN [score] >= 90 THEN 'A'
  WHEN [score] >= 80 THEN 'B'
  ELSE 'C'
END

-- running total, in Formula Multi-Row
SUM([amount]) OVER (ORDER BY [placed_at])

-- first non-null, in any expression field
coalesce([nickname], [first_name], 'Friend')
Data · Join → join 50-ROW SAMPLE
order_id customer amount
1043 Redwood Labs 1,284.00
1044 Bellhaven Co 312.50
1045 Orchard Group 9,940.10
1046 Marlow & Sons 76.20
1047 Kestrel Freight 2,410.00
1048 Juniper Analytics 587.35
04 · Test

A test run is 50 rows, and it shows you every one of them.

Press Test and the draft runs against a hard 50-row sample, fast enough to keep you in the loop, and it writes nothing to your real targets. Every port keeps its own data preview, so you can click any connection in the graph and see exactly what came out of it, with its schema and its messages.

There is no "run the whole thing to check" toggle, because that is the run you meant to schedule, not the one you meant to inspect.

05 · Ship

Draft, staging and production are different things, and the difference is enforced.

You edit a draft. Publishing puts a new version into staging, and only a promotion puts it into production. Editing a pipeline cannot break what is running in production, because what runs there is a version somebody promoted deliberately. Every run in history also resolves to the exact pipeline that produced it.

Always know where you stand. A chip on the builder shows unsaved edits, a draft ahead of staging, and which version staging and production each run.
1

Edit the draft

Test runs use it, on a 50-row sample that writes nothing, with a data preview on every port. Nothing you edit reaches staging or production until you say so.

2

Publish to staging

The draft is frozen as a new, numbered version, and staging starts running it. Nothing ever modifies that version again.

3

Promote to production

Production takes the version staging is running, and only a person allowed to reach production can do it. Every run is stamped with the version it executed.

4

Restore, any time

Any past version copies straight back into the draft and travels the same road. Rolling back is an edit, not an incident.

06 · Run

Four ways a pipeline starts, and one place they all end up.

On a schedule

A schedule belongs to staging or to production, fires in your instance's timezone, and runs whichever version its environment holds. Several servers never double-fire the same schedule.

When a file arrives

A watch trigger looks at a folder for files it has not seen before and starts a run when something new lands. The 2am partner drop gets handled without guessing at a cron time.

Fanned out across variants

A variant swaps the settings of chosen input and output nodes while the logic between them stays identical. Every run starts the base pipeline plus one run per variant: the same logic across regions, tenants or accounts.

As a step inside a Flow

A whole pipeline becomes one node in a bigger control-flow graph. Same queue, same history, same telemetry.

07 · Observe

Every run answers "what happened", not just "did it work".

The canvas animates node by node while a run progresses. Underneath, each node records its status, rows in, rows out, duration and error as it goes, and the totals are checked against the engine's own count when the run finishes.

Resource numbers, per run

Peak memory, peak spill and queue wait are recorded for every run. Memory answers "do I need a bigger server"; queue wait answers "do I need another one".

Cancellation that stops cleanly

A queued run cancels instantly. A running one is asked to stop and does so cleanly at the next node or batch boundary, without a killed process.

A dead server is noticed, and a job is never run twice

A running job checks in while it works. If its server goes quiet, the run is failed by name within minutes and reported. It is never left hanging, and never silently re-executed against your targets, because re-running a half-finished load is how rows get duplicated.

Rowsin and out, on every node
msduration, per node
Peakmemory and spill, per run
Logsand messages, per run
08 · Orchestrate

One pipeline is a graph of rows. A Flow is a graph of pipelines.

The layer that decides what runs after what, and what happens when something fails: sequence pipelines, run them in parallel, branch on success or failure, and alert someone by email, Slack or Discord, on a schedule or on demand.

nightly-close Production Run: Failed Run historyWed 02:00 · Failed Run Cancel
Load ordersRun pipeline · orders_to_warehouse
Succeeded
Load refundsRun pipeline · refunds_from_sftp
Succeeded
ReconcileRun pipeline · month_end_reconcile
Failed
Publish martsRun pipeline · publish_marts
Skipped
Alert the data teamNotify · #data-ops
Succeeded
Refresh dashboardsRun pipeline · refresh_dashboards
Skipped
Inside Load orders Succeeded
S3 Files Data Cleanse Snowflake

Every Run pipeline step is a whole pipeline like this one, run at the version the flow's environment holds.

#data-ops on Slack

No alerts yet tonight.

Trinium
Flow "nightly-close" — step "Alert the data team" fired.
Triggered by: "Reconcile" (Failed)

Run in order, or side by side

Connect two steps and the second waits for the first. Steps with nothing between them run at the same time.

Choose what happens next

✓ on success✕ on failure• always

When a step fails, everything that depends on it is skipped instead of running on bad data.

Tell the right people

A notify step emails a person or posts to Slack or Discord, naming the step that failed. A failed alert never fails the flow.

Schedule it, then watch it

Run it on a schedule or by hand, in staging or production, and watch each step light up. The run history keeps every past night.

09 · The engine

Rust from end to end, with a hard rule about memory.

The part of the product that decides whether a pipeline finishes at 4am or pages somebody.

On the Java virtual machine

Memory is safe, but a garbage collector manages it: runs pause while it works, and somebody has to size the heap and keep tuning it.

In C++

Fast, but keeping memory safe is left to the programmer, and memory mistakes are the commonest source of serious crashes and security flaws in C++ software.

In Rust, Trinium

As fast as C++, with memory safety checked by the compiler before the code ever runs. No collector, no pauses, and no heap to size.

It sizes itself

Trinium measures the machine it runs on, sets aside what the database and other services beside it need, and takes most of the rest as its memory budget. There is nothing to configure, though you can set the budget yourself.

Every node runs at once

A pipeline is a live stream, not a queue of steps: every node starts together and passes rows along in batches, so one pipeline uses every core, and a slow destination slows the source instead of filling memory. The compiler refuses code that could race, which is what makes that safe.

More data than memory

Sorts, joins and aggregations spill to disk when they outgrow the budget, and carry on. If the work still does not fit, the run stops with an error that names the node, not a crashed process and a run stuck on Running.

Tested by counting rows

A harness of 14 pipelines covering nearly every node type pushes a million rows through and checks that the totals come out exact. A lost or duplicated row fails the run and names the node.

How it is deployed

One binary

The web app, the API and the worker are one executable. A setting at startup decides which parts a server runs.

Postgres is the only bus

The queue, run progress, logs and cancellation all go through tables in one Postgres database, so there is no message broker to operate.

Scale by adding servers

Each server runs one pipeline at a time and uses all of it. Add a server to run another at the same moment; a bigger server never changes the price.

10 · Connect

Credentials live in one place, and never in a pipeline.

A connection is a named, configured instance of a connector. Nodes store only its identifier, so credentials never enter a pipeline manifest, a version snapshot, or anything you hand to a colleague.

Encrypted at rest, tested on demand

Secrets are sealed with authenticated encryption and never displayed again. Every connector has a Test button that makes a real connection.

Shared, granted, or private

Each shared connection carries an access list naming everyone on the server, a workspace, or one person, optionally with an expiry date.

TLS all the way down

Every database connector has an SSL mode you set per connection, from full certificate verification down to disabled.

See every connector

One name, a staging side and a production side

Every connection holds a staging side and, when you want one, a production side, behind a single name. The same pipeline reaches your staging database when it runs in staging and your production database when it runs in production.

A test run always uses the staging side, so building never needs production access.
Who may promote or run in production is decided per connection, and nobody holds it implicitly, administrators included.
Sharing a pipeline lets people see and edit it. Running it still needs its connections.
Python Script · node
# each input is already a dataframe
df = input_1
features = ["tenure", "spend", "tickets"]

# scikit-learn: one line in python/requirements.txt
from sklearn.linear_model import LogisticRegression
model = LogisticRegression().fit(df[features], df["churned"])

# assign the result — don't return it
df["score"] = model.predict_proba(df[features])[:, 1]
output_1 = df[df.score > 0.8]
11 · Escape hatch

When the visual tools run out, drop to Python.

Zero to five inputs and zero to five outputs, so the same node can be a source, a transform or a sink. pandas, polars, numpy and dateutil are ready to import, and you hand back whichever you like, or a plain list of rows. Anything else, scikit-learn or statsmodels included, is one line in your own python/requirements.txt.

Its own container, always

The interpreter runs in a separate sandboxed container with a memory cap, a read-only filesystem, no outbound network and no access to your database or encryption key. A worker cannot fork Python, because the type system forbids it.

Declare what you return

Declaring the output schema keeps every downstream node's column metadata alive with no run, which is exactly what custom-code nodes normally destroy.

Try it against your own data for 30 days.

Full Enterprise, every feature enabled, running on your own server in about ten minutes.