Transformers
Transformers are pipeline nodes that clean, reshape, or enrich tabular data. A runnable workflow starts with a Data Input, passes data through one or more transformers, and can finish with a Data Output.
Transformer Categories
The public catalog organizes transformers by purpose:
-
Combine & Shape: Merge or union data from multiple inputs, or reshape it with pivot and unpivot operations.
-
Clean & Format: Change text case, filter text with regular expressions, drop columns, remove duplicate rows, and clean text values.
-
Transform & Enhance: Rename, encode, normalize, or split columns, and fill missing values with Impute.
-
Detect & Manage: Detect or handle outliers, infer schema types, and change data types.
-
Organize: Sort rows by one or more columns.
Reusable Pipeline Steps
Transformers make data-cleaning and preparation rules reusable. A chain of transformers shows the processing order clearly and applies the same rules each time the pipeline runs.
Getting Started with Transformers
- Add and configure a Data Input node. Choose its dataset and sampling settings.
- Choose the transformations you need from the Transformer References.
- Add each transformer, connect the nodes, and configure its parameters.
- Add a Data Output node if you want to download a file or write the result to a configured cloud or Snowflake destination.
Use Chat Commands to ask the AI to build a pipeline, analyze a result, plan the work, or delegate a complex request.
The Transformer References describe each node's inputs, settings, and output.