Nodes & connectors

Everything a pipeline is built from.

The nodes your team drags onto the canvas, and the systems they read from and write to. It is the same list the builder shows.

Connectors

Every connector has a real driver and a Test button.

Configure one once, share it with the people who should have it, and reference it from any pipeline. Credentials are encrypted at rest and never enter a manifest.

Database

PostgreSQL
MySQL
SQL Server
MongoDB

Warehouse & query engines

Snowflake
BigQuery
Amazon Redshift
Databricks
ClickHouse
Trino / Starburst
Amazon Athena

Lakehouse

Apache Iceberg

Cloud & remote storage

Amazon S3
Azure Blob Storage
Google Cloud Storage
SFTP

SaaS & APIs

REST API
Google Sheets
Salesforce
HubSpot
Jira
Zendesk
SMTP (Email)

Browse the objects

Pick a catalog, schema and table from a live browser instead of typing a fully qualified name from memory.

Columns without a run

Most sources report their schema from catalog metadata. Snowflake does it without waking a warehouse, and ClickHouse does it even for a query you wrote by hand.

Explicit TLS

Every database connector exposes an SSL mode you choose, from verify-full down to disabled. Managed cloud databases verify against the bundled root store with nothing to configure.

Highlights

A few that do more than their name suggests.

Fuzzy Match

Finds records that mean the same thing despite typos, abbreviations or word order, and either clusters the duplicates or matches against a reference list. Catherine matches Kathryn; Smith matches Smyth.

Data Validation

Declare the rules and decide what a violation costs. Good rows carry on, and bad rows leave by their own output, each tagged with the rule it broke.

Data Compare

Checks one dataset against another and gives two answers: a summary of where they disagree, and the full row-by-row difference, missing and extra rows included.

Route

Sends each row down one of up to ten named branches, where the first matching rule wins and anything left over has a branch of its own. Every row lands exactly once.

API Lookup

Calls an API once per row and adds the answer as new columns: geocoding, address validation, exchange rates or your own service. Like a join, where the table is a web API.

Apache Iceberg

Reads and writes Iceberg tables through your Iceberg REST catalog: append or full refresh, create the table if it is missing, partitioned writes and time-travel reads. The write is the catalog commit, so there is no crawler to run afterwards.

Every node

Grouped the way the builder groups them.

Search by name or by what you need to do: pivot, dedupe and soap all find their node.

Showing everything

Input 30

Text Input

Type a small table straight onto the canvas: lookup values, constants, or a fixed set of rows, versioned with the pipeline.

CSV Files

Read rows from a CSV file and stream them into the pipeline.

JSON Files

Read array or newline-delimited JSON; nested objects become columns.

XML Files

Read one or more XML files, shredding a repeating element into typed rows.

Parquet Files

Read one or more columnar Apache Parquet files.

SQLite Files

Read a table, or the result of a SQL query, from a local SQLite database file.

Excel Files

Read one sheet, or every sheet, from an Excel workbook.

S3 Files

Read a single object or a whole prefix from an S3 bucket, then optionally archive, delete, or bookmark the files so a re-run doesn't reprocess them.

Azure Blob Files

Read a single blob or a whole prefix from an Azure Storage container, then optionally archive, delete, or bookmark the files so a re-run doesn't reprocess them.

GCS Files

Read a single object or a whole prefix from a Google Cloud Storage bucket, then optionally archive, delete, or bookmark the files so a re-run doesn't reprocess them.

SFTP Files

Read a single file or a whole folder from an SFTP server, then optionally archive, delete, or bookmark the files so a re-run doesn't reprocess them.

List Files

List the files in a folder as rows, with each one's path, name, size and modified time, without reading any of them.

Postgres

Read the result of a SQL query from PostgreSQL.

Trino / Starburst

Read a table or SQL query from a Trino / Starburst cluster.

Apache Iceberg

Read an Apache Iceberg table from a REST catalog, with time-travel and column pruning. Its columns show without a run.

MySQL

Read a table or SQL query from a MySQL database.

SQL Server

Read a table or SQL query from Microsoft SQL Server.

MongoDB

Read documents from a MongoDB collection into rows.

Snowflake

Read a table or SQL query from Snowflake. Picking a table shows its columns without a run, and without waking a warehouse.

BigQuery

Read a table or SQL query from BigQuery. Columns AND what the query will scan both show without a run.

Databricks

Read a table or SQL query from a Databricks SQL warehouse. Results stream back columnar.

Redshift

Read a table or SQL query from Amazon Redshift.

ClickHouse

Read a table or SQL query from ClickHouse. Columns show without a run, even for a query you write yourself.

Athena

Query data in S3 with Amazon Athena. Picking a table shows its columns without a run, because the catalog already knows them.

REST API

Read records from any HTTP/JSON API, with six pagination styles, retries, and typed columns straight out.

Salesforce

Read Salesforce objects with SOQL, paged and typed automatically.

HubSpot

Read HubSpot CRM objects: contacts, companies, deals and tickets.

Jira

Read Jira issues with JQL.

Zendesk

Read Zendesk tickets, users, or organizations.

Google Sheets

Read a tab and range from a Google spreadsheet. Column types are inferred unless you read everything as text.

Preparation 18

Filter

Split rows by a condition: matches go to True and the rest to False.

Route

Send each row down one of several named branches, where the first matching rule wins. For a single yes/no condition, use Filter.

Manage Columns

Keep, drop, reorder, rename, and retype columns.

Dynamic Select

Keep or drop columns by a rule: by type, by a name pattern, or by an expression.

Dynamic Rename

Rename every matching column at once. Add or strip a prefix, replace text, or fix the capitalisation.

Sort

Order rows by one or more columns.

Unique

Split rows into first-seen (Unique) and repeats (Duplicate) by key columns.

Formula

Add or replace columns with SQL expressions.

Formula Multi-Field

Apply one expression across several fields at once.

Formula Multi-Row

Compute a column from other rows with SQL window functions, such as running totals, ranking, lag/lead and moving averages.

Data Cleanse

Trim whitespace, remove characters, fix nulls, and standardise text.

Fill Missing

Fill missing values with a constant or a column statistic (mean/median/mode).

Generate Rows

Synthesize rows from nothing: a sequence of numbers, dates or date-times.

Record ID

Stamp each row with a sequential number or a UUID.

Sample

Keep a subset of rows: the first or last N, skip N, every Nth, a 1-in-N chance, or exactly N at random.

Data Validation

Check the data against rules. Rows that hold up leave Valid; the rest leave Invalid, tagged with what they broke.

Data Profile

Describe the data instead of changing it: one row per column, with nulls, distinct counts, ranges and the most common values.

Data Compare

Check one dataset against another and report every row, value and column type that disagrees.

Join 6

Join

Combine two inputs on matching keys, returning the matched rows plus the unmatched orphans from both sides.

Join Multiple

Combine 2 or more inputs into one on a shared key (or record position). Every row is kept by default, null-filled where they don't match.

Append Fields

Append every Source row's columns onto every Target row (a cross join).

Find & Replace

Look up an input field against a mapping table to replace found text or append matched fields.

Union

Stack the rows of 2 or more inputs into one, matched by field name (or record position), with the fields reconciled.

Fuzzy Match

Find records that mean the same thing despite typos, spelling, or word order. Cluster duplicates, or match against a reference list.

Parse 6

DateTime

Convert a column between real date/time values and formatted text.

RegEx

Match, extract, replace, or split text with a regular expression.

Text to Columns

Split a text field into multiple columns or rows by a delimiter.

JSON

Parse a JSON text column into typed columns, a name/value table, or auto-flattened columns.

URL

Split a URL column into its parts (scheme, host, domain, path, query, fragment), pull query parameters into columns, or explode them into rows.

XML

Parse an XML text column into typed columns, a name/value table, or auto-discovered columns.

Transform 3

Summarize

Group by one or more columns and compute aggregate values: sum, average, count, median and more.

Cross Tab

Pivot rows into a wide cross-tabulation.

Transpose

Unpivot wide columns into tall Name/Value rows (the inverse of Cross Tab).

Output 24

CSV File

Write the incoming rows to a CSV file.

JSON File

Write rows as an array or newline-delimited JSON.

Parquet File

Write the incoming rows to a columnar Parquet file.

SQLite File

Write the incoming rows into a table in a new SQLite database file.

Excel Report

Write a styled Excel report with a table, number formats and an optional chart.

S3 File

Write the pipeline's output to an S3 bucket with a templated key (dates, run id, variant) and a write mode.

Azure Blob File

Write the pipeline's output to an Azure Storage container with a templated blob path (dates, run id, variant) and a write mode.

GCS File

Write the pipeline's output to a Google Cloud Storage bucket with a templated key (dates, run id, variant) and a write mode.

SFTP File

Write the pipeline's output to an SFTP server with a templated path (dates, run id, variant) and a write mode.

Postgres

Load rows into a PostgreSQL table (append, truncate, or upsert).

Trino / Starburst

Load rows into a Trino / Starburst table (append or truncate; auto-creates).

MySQL

Load rows into a MySQL table.

SQL Server

Load rows into a SQL Server table.

MongoDB

Write rows into a MongoDB collection as documents.

Snowflake

Load rows into a Snowflake table, creating it if it doesn't exist.

BigQuery

Load rows into a BigQuery table with a load job, appending or replacing. Loading is free, and the table is created if needed.

Databricks

Load rows into a Databricks table (append or replace) through a Unity Catalog volume.

Redshift

Load rows into an Amazon Redshift table (append, replace, or upsert). Best for small and medium loads.

ClickHouse

Load rows into a ClickHouse table (append or replace; auto-creates a MergeTree).

REST API

Push rows to an HTTP API, one request per row or batched into arrays, with retries.

Send Email

Email the rows to people: one message per row, per group, or one for the whole run, with the data attached as CSV or Excel.

Salesforce

Write rows back to Salesforce: insert, update, upsert or delete, 200 records a request, with every rejected record reported.

Apache Iceberg

Write rows to an Apache Iceberg table through a REST catalog, appending or fully refreshing, with auto-create and partitioning. No crawler needed.

Google Sheets

Write rows to a tab in a Google spreadsheet, overwriting or appending, and optionally creating the tab.

Developer 4

API Lookup

Look up each row against an API and add the answer as new columns. It works like a JOIN where the table is a web API.

Python Script

Run your own Python over the data, for anything the built-in nodes don't cover.

NoteCanvas

Pin a free-text note to the canvas to document a pipeline.

ContainerCanvas

Group nodes, and switch a whole section off without breaking the flow.

Missing something you need?

Tell us during your trial. And if what you need is truly bespoke, the Python Script node gets you there without waiting for us.