Skip to main content

Encode Transform

Use Encode to turn categorical values into numerical features. Standard execution (Pandas) keeps each source column and adds encoded columns. Several methods produce a different schema in Big Data/AWS Glue execution (PySpark).

Basic Usage​

To encode categorical columns:

  1. Select the Encode transform from the transform menu.
  2. Select one or more fields under Target Columns.
  3. Choose a Method.
  4. Apply the transformation.
note

Target Columns lists categorical columns only.

If no categorical columns appear​

If Target Columns says that no categorical columns were found, none of the columns in the current input schema are identified as categorical.

Encode cannot find a categorical column

Add Infer Schema Type between the data source and Encode, then run or preview that upstream node. Reopen Encode to select the columns it identified as categorical under Target Columns.

Encode with Infer Schema Type connected upstream

Configuration Options​

Basic Options​

  • Target Columns: Select one or more categorical columns to encode.
  • Method: Choose one of these encodings:
    • One-Hot Encoding
    • Label Encoding
    • Ordinal Encoding
    • Binary Encoding
    • Frequency Encoding
    • Hashing Encoding
tip

Select the information icon beside Method to open the encoding method guide.

Standard Execution Output​

In Standard execution, encoded columns are added alongside their source columns.

MethodColumns appended for each selected source column
One-HotOne {source}_{category} column per observed non-null category
LabelOne {source}_label column
OrdinalOne {source}_ordinal column
BinaryA data-dependent number of bit columns, normally named {source}_0, {source}_1, and so on
FrequencyOne {source}_frequency column containing raw occurrence counts
HashingEight columns named {source}_hash_0 through {source}_hash_7

The Encode panel cannot remove source columns, set the order of ordinal categories, or change the Standard hashing width.

Encoding Methods​

One-Hot Encoding

Keeps the source column and adds one binary indicator for each observed category.

Use for: Unordered values such as red, blue, and green.

Label Encoding

Keeps the source column and assigns each category a unique integer. These codes may imply an order that does not exist.

Use for: Binary values or columns with many categories when the implied ordering is acceptable.

Ordinal Encoding

Keeps the source column and adds an integer-ranked column. Standard execution respects an existing ordered pandas categorical type. Otherwise, it infers an order that may not match the intended meaning.

Use for: Values with a defined order, such as shirt sizes.

caution

The Encode panel does not let you set category order. Do not assume S, M, L, and XL will receive that order unless it is already stored in the input column's categorical metadata.

Binary Encoding

Keeps the source column and adds binary-code bit columns. This usually creates fewer columns than one-hot encoding.

Use for: Columns with many categories where one-hot encoding would be too wide.

Frequency Encoding

Keeps the source column and adds each category's raw occurrence count.

Use for: Distinguishing rare and common values when frequency is meaningful.

Hashing Encoding

Maps categories through a hash function. Standard execution keeps the source and adds eight hash columns for each selected field.

Use for: High-cardinality fields where a fixed output width is useful.

The current Encode panel does not provide a hash-width control.

Execution Mode Differences​

MethodStandard execution (Pandas)Big Data/AWS Glue execution (PySpark)
One-HotRetains the source and appends one scalar column per category.Replaces each selected source with one {source}_encoded vector column.
LabelRetains the source and appends {source}_label.Replaces each selected source with {source}_indexed.
OrdinalRetains the source and appends {source}_ordinal; respects an existing ordered categorical dtype.Uses the same string-indexing behavior as Label and replaces the source with {source}_indexed.
BinaryRetains the source and appends bit columns.Not supported.
FrequencyRetains the source and appends {source}_frequency.Replaces each selected source with {source}_freq.
HashingRetains each source and appends eight scalar hash columns per source.Replaces the selected sources with one combined hashed_features vector, using 262,144 features by default.
Execution mode affects the schema

The same configuration can produce different column names, column counts, and data types when you switch between Standard and Big Data/AWS Glue.

Examples​

This example uses Standard execution.

Example: Encoding Product Categories

Input Dataset:

Product IDCategoryPrice
1Electronics500
2Clothing50
3Electronics750
4Home200
5Clothing75

Configuration:

  • Target Columns: Category
  • Method: One-Hot Encoding

Result:

Product IDCategoryPriceCategory_ClothingCategory_ElectronicsCategory_Home
1Electronics500010
2Clothing50100
3Electronics750010
4Home200001
5Clothing75100

Best Practices​

  1. Choose a method based on whether the categories are ordered and how many unique values they contain.
  2. Use Binary or Hashing Encoding if One-Hot would create too many columns.
  3. Use Ordinal Encoding for a semantic rank only when the input already has the correct ordered categorical metadata.
  4. In Standard execution, add Drop Columns later if downstream nodes should receive only the encoded fields.
  5. Keep the encoding scheme consistent between training and test data.

Troubleshooting​

  • A field is missing: If it does not appear under Target Columns, check that it is identified as categorical.
  • Ordinal Encoding produces an unexpected rank: Check whether the input column has ordered categorical metadata. You cannot change the order in the Encode panel.
  • Hashing Encoding produces an unsuitable output shape: Standard execution always creates eight columns per selected source in the current UI. Choose another method if needed.