Use cases

Seven jobs Trinium was built to take over.

Each one is drawn below as it appears on the canvas, from real nodes you can look up. Pick the one that sounds most like your week.

01 · Data-lake pipelines, without the glue

One job, spread across four services, with four consoles and four ways to fail.

A function to fetch it, a state machine to sequence it, a crawler to notice the schema, a job to transform it. Nobody designed that. It accreted.

In Trinium it is one graph you can see. An S3 Files input reads a whole prefix of CSV, JSON, Parquet, Excel or XML files, gzipped or zipped if need be, and a watch schedule starts a run within about a minute of new objects landing, remembering what it has read so each run picks up only what is new. Data Validation sends bad rows to a quarantine file through an output of their own instead of poisoning the load, and the result lands as Parquet in your curated zone, or as an Apache Iceberg commit through an Iceberg REST catalog.

The columns are known before you run, because Trinium reads them from the file header, the Parquet footer or the catalog itself. There is no crawler to schedule and nothing to go stale. When several pipelines have to happen in order, a Flow sequences them, branches on failure, and tells somebody.

It also speaks the layout your lake already uses. A folder named reading_date=2026-09-01 becomes a column, a filter on it skips the folders you did not ask for before anything is downloaded, and an output writes one folder per value and replaces only the folders a run produces. Paths take ${date} and ${date-1}, in your instance's timezone. All four cloud stores work this way: S3, Azure Blob, Google Cloud Storage and SFTP.

What collapses into what

  • A function per step → nodes on one canvas, all running concurrently
  • A state machine → a Flow, with the same sequencing and branching
  • A crawler on a schedule → schema read directly, at design time, with no run
  • Date-partitioned folders → read as a column, and written one folder per value
  • Four consoles → one run screen, per node, with the logs beside it
  • Six independent meters → $999 per server, the same number every month

Scope worth knowing. Trinium writes the files and commits Iceberg tables; it does not register tables in a cloud catalog, so Athena reads your curated zone through a table you define once.

02 · The Monday report, off the desktop

The report one person rebuilds by hand every week, on a licence only they have.

An export from the finance system, a spreadsheet of adjustments, some cleaning, a join, a summary, and twelve emails. Every analytics team has one, and it is usually the hardest to hand over.

On the canvas it is the same shape it always was. Excel Files and a SQL Server table come in, Join puts them together, Fuzzy Match collapses “ACME Corp” and “Acme Corporation” into one customer, and Summarize does the arithmetic. Every expression is ordinary SQL, and [bracket] column names still work, so muscle memory survives the move.

The result leaves two ways. Excel Report writes a styled workbook with table styles, number formats and a chart. Send Email runs in per-group mode, so each regional manager gets one message with only their own rows, attached as a plain workbook. Then it goes on a schedule and stops being anybody's Monday morning.

And because accounts cost nothing, the person who asks for a change can open it, run it at fifty rows and see what it does, instead of joining a queue.

It also stops living on a laptop. The database password moves out of the workflow into a shared connection, encrypted at rest, that only the people you choose can use, and every change to the pipeline is recorded with the name of whoever made it.

What it is built from

  • Excel Files: one sheet, a named sheet, or every sheet in the workbook
  • Join, Summarize, Cross Tab: the blending and the arithmetic
  • Fuzzy Match: one customer, however many ways it was spelled
  • Excel Report: table styles, number formats and an optional chart
  • Send Email: one message per run, per group, or per row
03 · The nightly warehouse load

Forty tables have to be in the warehouse by 6am, and somebody has to know if they aren't.

The load itself is not hard. What is hard is the part around it: knowing which of the forty failed, why, whether the ones after it ran anyway, and who found out.

Each source is a pipeline and the whole night is a Flow: a control-flow graph whose steps are entire pipeline runs. Edges are barriers, so independent loads run at the same time and dependent ones wait. An edge carries on success, on failure or always, so the failure branch is part of the design rather than a try/except somebody added later.

Every run writes per-node status, row counts, timings and errors while it is running, so at 6:04am you open one screen and see the node that failed, named as it is on your canvas rather than as a UUID. If a server dies mid-run, the run is failed by name within minutes and reported. It is never silently run again against your targets.

What it is built from

  • Warehouse input and output nodes: Snowflake, BigQuery, Databricks, Redshift, ClickHouse, Trino
  • Flows: sequence, parallel, and branch on outcome
  • Schedules: cron, or a run when a file lands in a watched bucket
  • Notification destinations: email, Slack or Discord on failure
  • Run telemetry: rows, duration, peak memory and queue wait, per run

Included in both editions. Orchestration is not a separate product here.

04 · Client & partner file onboarding

Every partner sends a slightly different file, and one of them reordered the columns.

Onboarding a new client's data eats integration time, and a column that quietly moved is how bad data gets in without anyone noticing.

Trinium reads a whole folder as one input, whether it holds CSV, JSON, XML, Parquet or Excel, over SFTP, S3, Azure Blob or GCS. There is one drift rule and it is the same rule for all of them: a file that shares at least one column name with the first is reconciled by name, with missing columns null-filled and extra columns dropped, and a file that shares none is skipped. Every deviation is named in the run's messages, so a partner who reordered their columns without telling you shows up as a line you can read rather than as values in the wrong fields.

Then Data Validation gates the rows: nine checks, and the failures leave through an output of their own, either to a quarantine table or straight back to the partner as an email with their own bad rows attached. A run can succeed and still tell you four files were skipped.

What it is built from

  • SFTP, S3, Azure and GCS file inputs: a folder, not a file, remembering what it has read so a run only picks up what is new
  • One schema-drift setting: reconcile, skip, or fail, on all nine multi-file inputs
  • Data Validation: valid rows one way, quarantine the other
  • Watch triggers: the run starts when the file lands, not on the hour
  • Gzip and zip decompressed on the way in, and one run per file when you want the isolation
05 · Reconciliation you can hand over

"The row counts match" is not the same as "the data is right".

Month-end close, a warehouse migration, a vendor swap, a rewritten pipeline: in each case the risk is never the rows that vanished. It is the eleven thousand that arrived with a decimal in the wrong place.

The Data Compare node takes two inputs, primary (the source of truth) and target (the thing you are checking), and gives you two outputs. One is a summary: one row per column, telling you where they disagree. The other is the full row-by-row diff, every cell that does not match, keyed so you can go and look at it. Missing rows and extra rows are both reported, and there is no cap on the diff, because a truncated report that says “Succeeded” is worse than no report.

It refuses, by name, to guess: a non-unique key and a key typed differently on the two sides are both errors rather than a best effort. Values compare as text, so a column that changed type reports the type once instead of drowning you in ten million value mismatches. Rounding or casting belongs upstream, in Formula or Data Cleanse, where you can see it.

The summary output is kept even on a full production run, so you can leave this wired into the nightly load permanently and read the answer the next morning.

Why this one earns its place

Reconciliation is usually somebody's weekend and a spreadsheet. Here it is a node with two inputs and two answers: one for the person who asks “are we clean?” and one for the person who has to go and fix it.

  • Two outputs: a per-column summary, and the row-level diff
  • Both missing and extra rows reported
  • A retyped column is named once, not ten million times
  • Ambiguous keys are a refusal, never a guess
06 · Retiring scripts and cron

The pipeline works. The problem is that one person understands it.

A folder of Python, a crontab, and a README that stopped being true two years ago. It runs fine, until the person who wrote it is on holiday and it doesn't.

On a canvas, the graph is the documentation: anyone can open it and see where the data comes from, what happens to it, and where it lands, without reading anything. The transformations are still real work, done with Filter, Formula, Join, Summarize and Data Cleanse, and every row expression is ordinary SQL rather than a language invented for the product.

What you gain that a script never had: an immutable version every time you publish, a run stamped with the version that produced it, any published version restorable into your draft, and a failure that emails somebody by name. What you keep: when a job is truly bespoke, the Python Script node runs your code in its own sandboxed container.

What replaces what

  • The crontab → a schedule you can see, edit and pause
  • git log, maybe → immutable published versions, restorable into the draft
  • print() and a log file → per-node rows, timings and errors on every run
  • "did it run?" → alerts on failure, to a person or a webhook
  • A .env nobody rotates → connections encrypted at rest, never in the pipeline
07 · Sending results back to the business

The warehouse knows the answer. The people who need it live somewhere else.

Account health, product usage and a churn score are computed nightly and sit in a table nobody in sales has ever opened.

The score does not have to come from somewhere else. A data scientist's own Python can be a step in the same pipeline: it receives the rows as a pandas or polars frame, hands back the scored table, and runs on the same schedule as the load instead of from a notebook.

Salesforce Output writes it back: insert, update, upsert or delete, two hundred records a request. It is a node of its own rather than a generic API call for one unglamorous reason: Salesforce answers HTTP 200 on a request in which individual records failed. A generic writer sees the 200 and reports success. This one reads the result for every record and tells you exactly which were rejected.

Not everything needs a CRM. The same answer can land in a shared Google Sheet, in any endpoint through REST API Output, or in somebody's inbox as an Excel attachment, which is where most people will read it anyway.

Where data goes back out

  • Salesforce Output: insert, update, upsert or soft delete, with per-record results
  • REST API Output: any endpoint, two body shapes
  • Google Sheets: overwrite or append a tab, and create it if it is missing
  • Send Email: the rows themselves, attached as CSV or Excel
  • API Lookup: per-row enrichment on the way through, as node configuration rather than a macro you build

Yours will be the eighth.

Take the pipeline that annoys you most and rebuild it during the trial. Thirty days, the full Enterprise product, on your own server. Ask us when you get stuck, and the people who built it will answer.