Data Input
Use Data Input to choose a workflow dataset and control the sample sent downstream. Sampling is always on in the current workflow UI and defaults to the first 100,000 rows.
Every root in a runnable workflow must be a configured Data Input. Add more than one when separate branches or a multi-input transform need different datasets.
Basic Usage
- Add Data Input to the canvas and open its configuration.
- Select an existing project dataset, or use Upload Data to create one. Choosing a different dataset saves it immediately and resets sampling to Head with 100,000 rows and streaming on.
- Review the sampling method and size. If you change either setting, select Apply.
- Connect the node to the first transformation in the workflow.
Changing the dataset or sample updates downstream dataset mappings and the data available to those nodes, then requests a new pipeline run.
Input Sources
You can bring data in from these sources:
- Existing dataset: Select data already available in the project.
- From Device: Upload a CSV (
.csv) or Excel (.xlsor.xlsx) file. - From PDF / Image: Extract tables from PDF, JPG, JPEG, PNG, TIFF, or TIF files and import the selected tables as CSV datasets. Files are limited to 50 MB, and PDFs are limited to 10 pages.
- From Web URL: Extract tables from a web page, including a JavaScript-rendered page, and import the selected tables as CSV datasets.
- Third Party Sources: If the feature is enabled for the deployment and the
current plan allows it, connect Amazon S3, Azure Blob Storage, Google Cloud
Storage, or Snowflake and sync data into the project. The three object-storage
browsers accept CSV and Excel (
.xlsor.xlsx) files.
Uploads and connections create project datasets. When an import finishes, select the new dataset in Data Input. Rhombus saves the selection immediately.
Sampling Configuration
Sampling and streaming are always on in the current workflow UI. Downstream transforms receive the sample, not the full dataset.
| Option | Current behavior |
|---|---|
| Sampling Method | Head takes the first rows, Systematic takes every nth row, and Reservoir creates a memory-efficient sample. |
| Sample Size (rows) | Sets an absolute row target and defaults to 100000. If both size and percentage are present, size takes priority. |
| Or Percentage (%) | Sets a target from 0.1 through 100 percent. Clear Sample Size to use the percentage. |
| Random Seed | Makes randomized samples reproducible. Of the current streaming-compatible methods, only Reservoir uses it. |
| Systematic Step | Sets the interval for Systematic sampling. If left empty, Rhombus calculates it from the requested sample size. |
You cannot turn off streaming in the workflow UI. Use Head, Systematic, or Reservoir. The menu also shows Random and Stratified, but you cannot apply the form with either option while streaming is on.
Column Names
When a dataset enters the pipeline, Rhombus replaces periods (.) in column
names with spaces. For example, customer.id becomes customer id. Use the
converted name in downstream node settings.
Troubleshooting
- Apply is unavailable: Select a dataset. A newly selected dataset is saved immediately; Apply stays unavailable until you change a sampling setting.
- A percentage does not change the row target: Clear Sample Size. An absolute size takes priority when both fields have values.
- Random or Stratified cannot be applied: Select Head, Systematic, or Reservoir, the methods supported by enforced streaming.
- A newly imported dataset is missing: Wait for the upload or connector sync to finish. Refresh the dataset list if necessary, then select the dataset.
See Data Output to export the transformed result.