Browse the objects
Pick a catalog, schema and table from a live browser instead of typing a fully qualified name from memory.
The nodes your team drags onto the canvas, and the systems they read from and write to. It is the same list the builder shows.
Configure one once, share it with the people who should have it, and reference it from any pipeline. Credentials are encrypted at rest and never enter a manifest.
Pick a catalog, schema and table from a live browser instead of typing a fully qualified name from memory.
Most sources report their schema from catalog metadata. Snowflake does it without waking a warehouse, and ClickHouse does it even for a query you wrote by hand.
Every database connector exposes an SSL mode you choose, from verify-full down to disabled. Managed cloud databases verify against the bundled root store with nothing to configure.
Finds records that mean the same thing despite typos, abbreviations or word order, and either clusters the duplicates or matches against a reference list. Catherine matches Kathryn; Smith matches Smyth.
Declare the rules and decide what a violation costs. Good rows carry on, and bad rows leave by their own output, each tagged with the rule it broke.
Checks one dataset against another and gives two answers: a summary of where they disagree, and the full row-by-row difference, missing and extra rows included.
Sends each row down one of up to ten named branches, where the first matching rule wins and anything left over has a branch of its own. Every row lands exactly once.
Calls an API once per row and adds the answer as new columns: geocoding, address validation, exchange rates or your own service. Like a join, where the table is a web API.
Reads and writes Iceberg tables through your Iceberg REST catalog: append or full refresh, create the table if it is missing, partitioned writes and time-travel reads. The write is the catalog commit, so there is no crawler to run afterwards.
Search by name or by what you need to do: pivot, dedupe and soap all find their node.
Showing everything
Type a small table straight onto the canvas: lookup values, constants, or a fixed set of rows, versioned with the pipeline.
Read rows from a CSV file and stream them into the pipeline.
Read array or newline-delimited JSON; nested objects become columns.
Read one or more XML files, shredding a repeating element into typed rows.
Read one or more columnar Apache Parquet files.
Read a table, or the result of a SQL query, from a local SQLite database file.
Read one sheet, or every sheet, from an Excel workbook.
Read a single object or a whole prefix from an S3 bucket, then optionally archive, delete, or bookmark the files so a re-run doesn't reprocess them.
Read a single blob or a whole prefix from an Azure Storage container, then optionally archive, delete, or bookmark the files so a re-run doesn't reprocess them.
Read a single object or a whole prefix from a Google Cloud Storage bucket, then optionally archive, delete, or bookmark the files so a re-run doesn't reprocess them.
Read a single file or a whole folder from an SFTP server, then optionally archive, delete, or bookmark the files so a re-run doesn't reprocess them.
List the files in a folder as rows, with each one's path, name, size and modified time, without reading any of them.
Read the result of a SQL query from PostgreSQL.
Read a table or SQL query from a Trino / Starburst cluster.
Read an Apache Iceberg table from a REST catalog, with time-travel and column pruning. Its columns show without a run.
Read a table or SQL query from a MySQL database.
Read a table or SQL query from Microsoft SQL Server.
Read documents from a MongoDB collection into rows.
Read a table or SQL query from Snowflake. Picking a table shows its columns without a run, and without waking a warehouse.
Read a table or SQL query from BigQuery. Columns AND what the query will scan both show without a run.
Read a table or SQL query from a Databricks SQL warehouse. Results stream back columnar.
Read a table or SQL query from Amazon Redshift.
Read a table or SQL query from ClickHouse. Columns show without a run, even for a query you write yourself.
Query data in S3 with Amazon Athena. Picking a table shows its columns without a run, because the catalog already knows them.
Read records from any HTTP/JSON API, with six pagination styles, retries, and typed columns straight out.
Read Salesforce objects with SOQL, paged and typed automatically.
Read HubSpot CRM objects: contacts, companies, deals and tickets.
Read Jira issues with JQL.
Read Zendesk tickets, users, or organizations.
Read a tab and range from a Google spreadsheet. Column types are inferred unless you read everything as text.
Split rows by a condition: matches go to True and the rest to False.
Send each row down one of several named branches, where the first matching rule wins. For a single yes/no condition, use Filter.
Keep, drop, reorder, rename, and retype columns.
Keep or drop columns by a rule: by type, by a name pattern, or by an expression.
Rename every matching column at once. Add or strip a prefix, replace text, or fix the capitalisation.
Order rows by one or more columns.
Split rows into first-seen (Unique) and repeats (Duplicate) by key columns.
Add or replace columns with SQL expressions.
Apply one expression across several fields at once.
Compute a column from other rows with SQL window functions, such as running totals, ranking, lag/lead and moving averages.
Trim whitespace, remove characters, fix nulls, and standardise text.
Fill missing values with a constant or a column statistic (mean/median/mode).
Synthesize rows from nothing: a sequence of numbers, dates or date-times.
Stamp each row with a sequential number or a UUID.
Keep a subset of rows: the first or last N, skip N, every Nth, a 1-in-N chance, or exactly N at random.
Check the data against rules. Rows that hold up leave Valid; the rest leave Invalid, tagged with what they broke.
Describe the data instead of changing it: one row per column, with nulls, distinct counts, ranges and the most common values.
Check one dataset against another and report every row, value and column type that disagrees.
Combine two inputs on matching keys, returning the matched rows plus the unmatched orphans from both sides.
Combine 2 or more inputs into one on a shared key (or record position). Every row is kept by default, null-filled where they don't match.
Append every Source row's columns onto every Target row (a cross join).
Look up an input field against a mapping table to replace found text or append matched fields.
Stack the rows of 2 or more inputs into one, matched by field name (or record position), with the fields reconciled.
Find records that mean the same thing despite typos, spelling, or word order. Cluster duplicates, or match against a reference list.
Convert a column between real date/time values and formatted text.
Match, extract, replace, or split text with a regular expression.
Split a text field into multiple columns or rows by a delimiter.
Parse a JSON text column into typed columns, a name/value table, or auto-flattened columns.
Split a URL column into its parts (scheme, host, domain, path, query, fragment), pull query parameters into columns, or explode them into rows.
Parse an XML text column into typed columns, a name/value table, or auto-discovered columns.
Group by one or more columns and compute aggregate values: sum, average, count, median and more.
Pivot rows into a wide cross-tabulation.
Unpivot wide columns into tall Name/Value rows (the inverse of Cross Tab).
Write the incoming rows to a CSV file.
Write rows as an array or newline-delimited JSON.
Write the incoming rows to a columnar Parquet file.
Write the incoming rows into a table in a new SQLite database file.
Write a styled Excel report with a table, number formats and an optional chart.
Write the pipeline's output to an S3 bucket with a templated key (dates, run id, variant) and a write mode.
Write the pipeline's output to an Azure Storage container with a templated blob path (dates, run id, variant) and a write mode.
Write the pipeline's output to a Google Cloud Storage bucket with a templated key (dates, run id, variant) and a write mode.
Write the pipeline's output to an SFTP server with a templated path (dates, run id, variant) and a write mode.
Load rows into a PostgreSQL table (append, truncate, or upsert).
Load rows into a Trino / Starburst table (append or truncate; auto-creates).
Load rows into a MySQL table.
Load rows into a SQL Server table.
Write rows into a MongoDB collection as documents.
Load rows into a Snowflake table, creating it if it doesn't exist.
Load rows into a BigQuery table with a load job, appending or replacing. Loading is free, and the table is created if needed.
Load rows into a Databricks table (append or replace) through a Unity Catalog volume.
Load rows into an Amazon Redshift table (append, replace, or upsert). Best for small and medium loads.
Load rows into a ClickHouse table (append or replace; auto-creates a MergeTree).
Push rows to an HTTP API, one request per row or batched into arrays, with retries.
Email the rows to people: one message per row, per group, or one for the whole run, with the data attached as CSV or Excel.
Write rows back to Salesforce: insert, update, upsert or delete, 200 records a request, with every rejected record reported.
Write rows to an Apache Iceberg table through a REST catalog, appending or fully refreshing, with auto-create and partitioning. No crawler needed.
Write rows to a tab in a Google spreadsheet, overwriting or appending, and optionally creating the tab.
Look up each row against an API and add the answer as new columns. It works like a JOIN where the table is a web API.
Run your own Python over the data, for anything the built-in nodes don't cover.
Pin a free-text note to the canvas to document a pipeline.
Group nodes, and switch a whole section off without breaking the flow.
No nodes match that.
Try a broader word, or ask us. If it is a real gap, we would rather know.
Tell us during your trial. And if what you need is truly bespoke, the Python Script node gets you there without waiting for us.