You need to get a 12 GB CSV into Parquet for Athena. You reach for pandas, write pd.read_csv('huge.csv').to_parquet('out.parquet') — and watch your process get killed on an out-of-memory error. Or you scaffold yet another throwaway script that you'll never find again.
DataMorph is a single CLI command that converts between data formats by streaming row-by-row, so files over 10 GB fit in constant memory. It infers schemas automatically, batch-converts whole directories, and validates data against a schema in CI with proper exit codes. No DataFrame boilerplate, no OOM.
pandas is fantastic, but it loads the entire file into memory. On anything past a few hundred MB that means swap, then an OOM kill. DataMorph streams:
--recursiveDataMorph is not on public PyPI. Install directly from GitHub (recommended), or via Homebrew/Scoop:
# pip (recommended)
pip install git+https://github.com/Coding-Dev-Tools/datamorph.git
# Homebrew (macOS/Linux)
brew tap Coding-Dev-Tools/tap
brew install datamorph
# Scoop (Windows)
scoop bucket add Coding-Dev-Tools https://github.com/Coding-Dev-Tools/scoop-bucket
scoop install datamorph
Then verify it works:
datamorph --help
The core command is convert — source path in, destination path out. The output format is detected from the file extension:
# CSV → Parquet (the classic "convert csv to parquet cli" task)
datamorph convert input.csv output.parquet
# JSON → CSV for reporting
datamorph convert input.json output.csv
# YAML → JSON
datamorph convert input.yaml output.json
# Parquet → CSV for a quick look
datamorph convert input.parquet output.csv
No flags, no config. DataMorph reads the schema from the data and streams the conversion. This is the command you'll actually type when someone drops a 9 GB export on your desk.
One file at a time is fine; an entire folder of mixed files is where batch mode earns its keep:
# Convert every CSV in ./csv_data/ to Parquet in ./parquet_data/, recursively
datamorph batch ./csv_data/ ./parquet_data/ --from csv --to parquet --recursive
batch walks the tree, converts each matching file, and preserves the directory structure. Point it at a raw exports folder and get a processed folder out the other side.
Before you depend on a conversion, you often want to know what the data looks like. schema infers the structure and can emit it as JSON:
# Pretty-print the inferred schema
datamorph schema data.parquet
# Export schema for later validation
datamorph schema data.csv --json-output > schema.json
Use the exported schema as a contract: validate future data files against it so a silent column change never reaches production.
The highest-leverage use of DataMorph is a CI gate. validate exits with code 1 on a schema mismatch — perfect for blocking a bad data file before it ships:
# GitHub Actions — fail the build on schema drift
- name: Validate data schemas
run: |
pip install git+https://github.com/Coding-Dev-Tools/datamorph.git
datamorph validate data/events.csv --schema schemas/events.json --strict
datamorph validate data/users.json --schema schemas/users.json --strict
# Generic CI — fail on schema mismatch
datamorph validate data.csv --schema schema.json || echo "Schema mismatch detected!"
--strict tightens the check (unknown fields and type drift become failures). --json-output gives you machine-readable results for dashboards or audit trails.
Every data team eventually writes conversion glue. Here's how a dedicated CLI compares:
| Capability | pandas script | csvkit | DataMorph |
|---|---|---|---|
| CSV → Parquet | ✅ | ❌ | ✅ |
| Streaming >10 GB (no OOM) | ❌ (loads all) | ❌ | ✅ |
| Schema validation + exit codes | ❌ (you write it) | ❌ | ✅ |
| Batch directory convert | ❌ (write a loop) | ❌ | ✅ --recursive |
| 6+ format pairs (incl. Avro/Parquet) | ✅ (with libs) | ❌ (CSV only) | ✅ |
| Zero-config CLI | ❌ (imports + boilerplate) | ✅ | ✅ |
DataMorph vs pandas: pandas loads the whole file into a DataFrame. DataMorph streams row-by-row — no OOM on large files and no import boilerplate.
DataMorph vs csvkit: csvkit only handles CSV. DataMorph covers 6+ formats including Parquet and Avro.
DataMorph is one of 11 tools in the DevForge suite. One license covers all CLI tools.
| Plan | Price | Best For |
|---|---|---|
| Free | $0 | Individuals, OSS — CLI only, 100 conversions/month |
| DataMorph Pro | $12/mo | Professionals — unlimited conversions, streaming, batch mode, all formats |
| Suite (all 11 tools) | $49/mo ($39/mo annual) | Full toolkit — 40% savings |
🔹 No lock-in: the CLI works fully offline on the free tier — no telemetry, no phone-home.
datamorph convert input.csv output.parquet streams it--recursive) converts entire directories in a single callvalidate command exits 1 on schema mismatch — drop it into CI to block bad data before deploymentAll claims verified against datamorph/README.md (live on-disk). Install via pip install git+https://github.com/Coding-Dev-Tools/datamorph.git; DataMorph is not on public PyPI (PyPI JSON API returns 404 for datamorph-cli, verified this run), so the git+ form is the working install. Homebrew/Scoop formulas verified in homebrew-tap/ and scoop-bucket/. Formats: CSV, JSON, JSONL, YAML, Parquet, Avro, Protobuf. Commands: convert, batch, schema, validate, formats. Pricing: Free $0 / 100 conversions per month, Pro $12/mo (unlimited + streaming + batch + all formats), Suite $49/mo ($39 annual, all 11 tools). License: MIT. Python 3.10+.