Search every node
One search box covers every node, matching name, category and description. It ignores which tab you are on, so you never have to know where a node lives.
Build on a canvas, see the columns before anything runs, test on fifty rows, then publish to staging and promote to production. Here is each step, and the engine underneath all of them.
Drag a node out of the ribbon, drop it on the canvas and wire it up. Each node opens a configuration panel written specifically for it, with the real columns from upstream already in the pickers. You will not find a generic key/value editor anywhere.
One search box covers every node, matching name, category and description. It ignores which tab you are on, so you never have to know where a node lives.
Every node carries its own documentation: what it does, when to reach for it, worked examples and the gotchas. You can read it from the ribbon before you drop the node and from the config panel afterwards.
Pin notes to the canvas, and group nodes into a Container you can switch off in one click. The data then passes straight through the section as if it were not there.
Trinium walks your graph and propagates types the same way the executor propagates data, without moving a single row. Parquet footers, CSV headers, database catalogs and warehouse metadata answer the question for most sources before anything runs.
Every field that takes a column name offers the real ones, typed, the moment the node is wired.
A formula is validated against the real inbound schema. A misspelled column or a type mismatch shows up in the builder instead of as a failed run at 3am.
If one node cannot be known without running, only that branch goes quiet. The rest of the canvas stays lit.
Filter conditions, Formula columns, multi-field rewrites, window functions and validation rules all use the same syntax. There is no proprietary dialect to learn and no second language hiding in a different node.
Write [unit price] if that is your muscle memory,
or "unit price", or plain
price. All three resolve to the same column.
Logical, text, number, date and conversion functions, each one documented and tested. Browse them inside the expression field or on a help page in the app.
Running totals, lag and lead, moving averages, ranking and tiling are all standard SQL, so there is no special node vocabulary to memorise.
-- graded, with a fallback, in one expression CASE WHEN [score] >= 90 THEN 'A' WHEN [score] >= 80 THEN 'B' ELSE 'C' END -- running total, in Formula Multi-Row SUM([amount]) OVER (ORDER BY [placed_at]) -- first non-null, in any expression field coalesce([nickname], [first_name], 'Friend')
| order_id | customer | amount |
|---|---|---|
| 1043 | Redwood Labs | 1,284.00 |
| 1044 | Bellhaven Co | 312.50 |
| 1045 | Orchard Group | 9,940.10 |
| 1046 | Marlow & Sons | 76.20 |
| 1047 | Kestrel Freight | 2,410.00 |
| 1048 | Juniper Analytics | 587.35 |
Press Test and the draft runs against a hard 50-row sample, fast enough to keep you in the loop, and it writes nothing to your real targets. Every port keeps its own data preview, so you can click any connection in the graph and see exactly what came out of it, with its schema and its messages.
There is no "run the whole thing to check" toggle, because that is the run you meant to schedule, not the one you meant to inspect.
You edit a draft. Publishing puts a new version into staging, and only a promotion puts it into production. Editing a pipeline cannot break what is running in production, because what runs there is a version somebody promoted deliberately. Every run in history also resolves to the exact pipeline that produced it.
Test runs use it, on a 50-row sample that writes nothing, with a data preview on every port. Nothing you edit reaches staging or production until you say so.
The draft is frozen as a new, numbered version, and staging starts running it. Nothing ever modifies that version again.
Production takes the version staging is running, and only a person allowed to reach production can do it. Every run is stamped with the version it executed.
Any past version copies straight back into the draft and travels the same road. Rolling back is an edit, not an incident.
A schedule belongs to staging or to production, fires in your instance's timezone, and runs whichever version its environment holds. Several servers never double-fire the same schedule.
A watch trigger looks at a folder for files it has not seen before and starts a run when something new lands. The 2am partner drop gets handled without guessing at a cron time.
A variant swaps the settings of chosen input and output nodes while the logic between them stays identical. Every run starts the base pipeline plus one run per variant: the same logic across regions, tenants or accounts.
A whole pipeline becomes one node in a bigger control-flow graph. Same queue, same history, same telemetry.
The canvas animates node by node while a run progresses. Underneath, each node records its status, rows in, rows out, duration and error as it goes, and the totals are checked against the engine's own count when the run finishes.
Peak memory, peak spill and queue wait are recorded for every run. Memory answers "do I need a bigger server"; queue wait answers "do I need another one".
A queued run cancels instantly. A running one is asked to stop and does so cleanly at the next node or batch boundary, without a killed process.
A running job checks in while it works. If its server goes quiet, the run is failed by name within minutes and reported. It is never left hanging, and never silently re-executed against your targets, because re-running a half-finished load is how rows get duplicated.
The layer that decides what runs after what, and what happens when something fails: sequence pipelines, run them in parallel, branch on success or failure, and alert someone by email, Slack or Discord, on a schedule or on demand.
Every Run pipeline step is a whole pipeline like this one, run at the version the flow's environment holds.
No alerts yet tonight.
Flow "nightly-close" — step "Alert the data team" fired. Triggered by: "Reconcile" (Failed)
Connect two steps and the second waits for the first. Steps with nothing between them run at the same time.
✓ on success✕ on failure• always
When a step fails, everything that depends on it is skipped instead of running on bad data.
A notify step emails a person or posts to Slack or Discord, naming the step that failed. A failed alert never fails the flow.
Run it on a schedule or by hand, in staging or production, and watch each step light up. The run history keeps every past night.
The part of the product that decides whether a pipeline finishes at 4am or pages somebody.
Memory is safe, but a garbage collector manages it: runs pause while it works, and somebody has to size the heap and keep tuning it.
Fast, but keeping memory safe is left to the programmer, and memory mistakes are the commonest source of serious crashes and security flaws in C++ software.
As fast as C++, with memory safety checked by the compiler before the code ever runs. No collector, no pauses, and no heap to size.
Trinium measures the machine it runs on, sets aside what the database and other services beside it need, and takes most of the rest as its memory budget. There is nothing to configure, though you can set the budget yourself.
A pipeline is a live stream, not a queue of steps: every node starts together and passes rows along in batches, so one pipeline uses every core, and a slow destination slows the source instead of filling memory. The compiler refuses code that could race, which is what makes that safe.
Sorts, joins and aggregations spill to disk when they outgrow the budget, and carry on. If the work still does not fit, the run stops with an error that names the node, not a crashed process and a run stuck on Running.
A harness of 14 pipelines covering nearly every node type pushes a million rows through and checks that the totals come out exact. A lost or duplicated row fails the run and names the node.
The web app, the API and the worker are one executable. A setting at startup decides which parts a server runs.
The queue, run progress, logs and cancellation all go through tables in one Postgres database, so there is no message broker to operate.
Each server runs one pipeline at a time and uses all of it. Add a server to run another at the same moment; a bigger server never changes the price.
A connection is a named, configured instance of a connector. Nodes store only its identifier, so credentials never enter a pipeline manifest, a version snapshot, or anything you hand to a colleague.
Secrets are sealed with authenticated encryption and never displayed again. Every connector has a Test button that makes a real connection.
Each shared connection carries an access list naming everyone on the server, a workspace, or one person, optionally with an expiry date.
Every database connector has an SSL mode you set per connection, from full certificate verification down to disabled.
Every connection holds a staging side and, when you want one, a production side, behind a single name. The same pipeline reaches your staging database when it runs in staging and your production database when it runs in production.
# each input is already a dataframe df = input_1 features = ["tenure", "spend", "tickets"] # scikit-learn: one line in python/requirements.txt from sklearn.linear_model import LogisticRegression model = LogisticRegression().fit(df[features], df["churned"]) # assign the result — don't return it df["score"] = model.predict_proba(df[features])[:, 1] output_1 = df[df.score > 0.8]
Zero to five inputs and zero to five outputs, so the same node can be a source, a
transform or a sink. pandas, polars, numpy and dateutil are ready to import, and you
hand back whichever you like, or a plain list of rows. Anything else, scikit-learn
or statsmodels included, is one line in your own
python/requirements.txt.
The interpreter runs in a separate sandboxed container with a memory cap, a read-only filesystem, no outbound network and no access to your database or encryption key. A worker cannot fork Python, because the type system forbids it.
Declaring the output schema keeps every downstream node's column metadata alive with no run, which is exactly what custom-code nodes normally destroy.
Full Enterprise, every feature enabled, running on your own server in about ten minutes.