← Back to Blog

Convert CSV to Parquet from the command line with DataMorph

Product: DataMorph · Category: Data Engineering, ETL, CLI · Tags: csv to parquet data format converter streaming etl schema validation

You need to get a 12 GB CSV into Parquet for Athena. You reach for pandas, write pd.read_csv('huge.csv').to_parquet('out.parquet') — and watch your process get killed on an out-of-memory error. Or you scaffold yet another throwaway script that you'll never find again.

DataMorph is a single CLI command that converts between data formats by streaming row-by-row, so files over 10 GB fit in constant memory. It infers schemas automatically, batch-converts whole directories, and validates data against a schema in CI with proper exit codes. No DataFrame boilerplate, no OOM.

Why a dedicated converter instead of pandas?

pandas is fantastic, but it loads the entire file into memory. On anything past a few hundred MB that means swap, then an OOM kill. DataMorph streams:

Install DataMorph

DataMorph is not on public PyPI. Install directly from GitHub (recommended), or via Homebrew/Scoop:

# pip (recommended)
pip install git+https://github.com/Coding-Dev-Tools/datamorph.git

# Homebrew (macOS/Linux)
brew tap Coding-Dev-Tools/tap
brew install datamorph

# Scoop (Windows)
scoop bucket add Coding-Dev-Tools https://github.com/Coding-Dev-Tools/scoop-bucket
scoop install datamorph

Then verify it works:

datamorph --help

Quick start: convert a file in one command

The core command is convert — source path in, destination path out. The output format is detected from the file extension:

# CSV → Parquet (the classic "convert csv to parquet cli" task)
datamorph convert input.csv output.parquet

# JSON → CSV for reporting
datamorph convert input.json output.csv

# YAML → JSON
datamorph convert input.yaml output.json

# Parquet → CSV for a quick look
datamorph convert input.parquet output.csv

No flags, no config. DataMorph reads the schema from the data and streams the conversion. This is the command you'll actually type when someone drops a 9 GB export on your desk.

Batch-convert whole directories

One file at a time is fine; an entire folder of mixed files is where batch mode earns its keep:

# Convert every CSV in ./csv_data/ to Parquet in ./parquet_data/, recursively
datamorph batch ./csv_data/ ./parquet_data/ --from csv --to parquet --recursive

batch walks the tree, converts each matching file, and preserves the directory structure. Point it at a raw exports folder and get a processed folder out the other side.

Inspect and export schemas

Before you depend on a conversion, you often want to know what the data looks like. schema infers the structure and can emit it as JSON:

# Pretty-print the inferred schema
datamorph schema data.parquet

# Export schema for later validation
datamorph schema data.csv --json-output > schema.json

Use the exported schema as a contract: validate future data files against it so a silent column change never reaches production.

Validate data in CI (the real ROI)

The highest-leverage use of DataMorph is a CI gate. validate exits with code 1 on a schema mismatch — perfect for blocking a bad data file before it ships:

# GitHub Actions — fail the build on schema drift
- name: Validate data schemas
  run: |
    pip install git+https://github.com/Coding-Dev-Tools/datamorph.git
    datamorph validate data/events.csv --schema schemas/events.json --strict
    datamorph validate data/users.json --schema schemas/users.json --strict
# Generic CI — fail on schema mismatch
datamorph validate data.csv --schema schema.json || echo "Schema mismatch detected!"

--strict tightens the check (unknown fields and type drift become failures). --json-output gives you machine-readable results for dashboards or audit trails.

Comparison: DataMorph vs the usual approaches

Every data team eventually writes conversion glue. Here's how a dedicated CLI compares:

Capabilitypandas scriptcsvkitDataMorph
CSV → Parquet✅❌✅
Streaming >10 GB (no OOM)❌ (loads all)❌✅
Schema validation + exit codes❌ (you write it)❌✅
Batch directory convert❌ (write a loop)❌✅ --recursive
6+ format pairs (incl. Avro/Parquet)✅ (with libs)❌ (CSV only)✅
Zero-config CLI❌ (imports + boilerplate)✅✅

DataMorph vs pandas: pandas loads the whole file into a DataFrame. DataMorph streams row-by-row — no OOM on large files and no import boilerplate.

DataMorph vs csvkit: csvkit only handles CSV. DataMorph covers 6+ formats including Parquet and Avro.

Pricing

DataMorph is one of 11 tools in the DevForge suite. One license covers all CLI tools.

PlanPriceBest For
Free$0Individuals, OSS — CLI only, 100 conversions/month
DataMorph Pro$12/moProfessionals — unlimited conversions, streaming, batch mode, all formats
Suite (all 11 tools)$49/mo ($39/mo annual)Full toolkit — 40% savings

🔹 No lock-in: the CLI works fully offline on the free tier — no telemetry, no phone-home.

Key takeaways

Next steps

Verification notes

All claims verified against datamorph/README.md (live on-disk). Install via pip install git+https://github.com/Coding-Dev-Tools/datamorph.git; DataMorph is not on public PyPI (PyPI JSON API returns 404 for datamorph-cli, verified this run), so the git+ form is the working install. Homebrew/Scoop formulas verified in homebrew-tap/ and scoop-bucket/. Formats: CSV, JSON, JSONL, YAML, Parquet, Avro, Protobuf. Commands: convert, batch, schema, validate, formats. Pricing: Free $0 / 100 conversions per month, Pro $12/mo (unlimited + streaming + batch + all formats), Suite $49/mo ($39 annual, all 11 tools). License: MIT. Python 3.10+.