Convert CSV ↔ Parquet ↔ DuckDB ↔ zip-CSV ¶
When to use this ¶
You have data in format X (a regulator handed you a folder of CSVs, an upstream pipeline emits parquet, a teammate sent a DuckDB file) and a downstream tool wants format Y. Or you’re loading the same package over and over and want to switch to a format with faster cold-start.
Quick example ¶
Convert the bundled Leavenworth CSV directory into a parquet directory. Destination extension determines the output format:
Expected:
The same network reloads ~5× faster than the CSV directory it came from.
Step-by-step ¶
1. Pick the destination format ¶
Each format has different trade-offs for cold-read latency, write speed, and predicate pushdown. Rule of thumb: parquet for interchange, duckdb for repeated local use, zipcsv for emailing, csv only when a downstream tool demands it.
| Format | Cold read | Write | Predicate pushdown | Portability |
|---|---|---|---|---|
csv (directory) |
Slowest — full scan, no schema. | Slow — text encoding. | None. | Universal. Any tool reads CSV. |
parquet (directory or file) |
Fast — columnar, schema embedded. | Fast. | Yes — column + row-group prune. | Wide. Arrow, DuckDB, Polars, pandas. |
duckdb (single file) |
Fastest — pre-indexed, statistics cached. | Slowest — DDL + insert per table. | Yes — full SQL. | DuckDB-only readers. |
zipcsv (single .zip) |
Slow — text decode after unzip. | Medium. | None. | Universal + single-file. |
2. Run netstead convert ¶
The CLI infers format from the destination extension (.parquet, .duckdb, .zip) or from whether the target is an existing directory:
Override the inference with --format when the extension is missing or ambiguous:
netstead convert ./csv_dir ./out.duckdb --format duckdb
netstead convert ./csv_dir ./out --format parquet
The same dispatch is available from Python — Network.write and Package.write accept the same destination + optional format= kwarg as the CLI:
import tempfile, pathlib
from netstead import Network
from netstead.fixtures import leavenworth
net = Network.from_source(leavenworth.csv_dir())
with tempfile.TemporaryDirectory() as d:
net.write(pathlib.Path(d) / "leavenworth.duckdb")
3. (Optional) Pick an engine ¶
Switch to the pandas engine if you hit the null-typed column issue in #163 — DuckDB’s strict typing rejects columns that are entirely null in the source, while pandas coerces them to object/None:
4. Verify the round-trip ¶
Load the destination back and validate against the spec. This catches any schema drift introduced by the conversion (e.g. a string column that came back as int because every value happened to parse):
import tempfile, pathlib
from corral.reports import Severity
from netstead import Network
from netstead.fixtures import leavenworth
# Round-trip CSV → parquet → re-load → validate.
with tempfile.TemporaryDirectory() as d:
out = pathlib.Path(d) / "leavenworth.parquet"
Network.from_source(leavenworth.csv_dir()).write(out)
net = Network.from_source(out)
report = net.validate()
assert not report.has_errors, [i.code for i in report.issues if i.severity is Severity.ERROR]
print(f"{net.spec_version}: {net.links.count()} links — round-trip clean")
See validate-network for what’s in the report.
5. (Optional) Convert remote → local ¶
convert accepts the same URL surface as from_source, so cloud → local is a one-liner. The download happens once; from then on you load the local copy:
See Read from S3 for credential handling.
Common variations ¶
Default — CLI conversion with extension inference
Most conversions are one-line CLI calls. The destination extension picks the format.
Programmatic conversion in Python
Same dispatch as the CLI; useful inside scripts and notebooks.
Re-pack only a subset of tables
Keep just the tables you need; the writer drops the rest.
See the scope recipes for FK-aware variants that walk relationships.
Partitioned parquet (one file per table)
Pass a directory destination with --format parquet. The writer creates one file per table; per-table row-group partitioning is the parquet engine’s default.
Compress the CSV output (zipcsv)
Same wire format as CSV but typically 5–10× smaller — useful for emailing or storing in artifact stores.
Validate during the write
Raises before writing if any ERROR finding fires.
Pitfalls ¶
- Null-typed columns can fail strict-write backends. If a column is 100% null in the source CSV, DuckDB may reject the write. Workaround: open the source, fill the column with a typed default, and write. Tracked in #163.
- Extension inference is case-sensitive.
out.PARQUETwon’t be recognised as parquet — pass--format parquet. - Converting to CSV loses dtypes. Re-reading the CSV without a schema (no GMNS spec match) infers columns afresh. Round-trip safety relies on the schema being re-applied at load time, which happens automatically for GMNS-shaped directories.
- DuckDB files lock per-process. Two Python processes can’t open the same
.duckdbfile in write mode at once. For shared use, write parquet instead; for one-writer-many-readers, ensure the writer closes the connection (the Engine does this when theNetworkis garbage-collected ornet.close()is called). - Zip-CSV holds the whole archive in memory on read. Fine for Leavenworth-scale networks; not fine for regional ones. Convert zip-CSV to parquet up-front if you’ll re-read.
- Overwriting an existing destination is allowed.
convertwon’t ask before clobbering — chain atest -e dest && exit 1in CI scripts if you need an explicit guard.
See also ¶
- Read from S3 —
convertworks across cloud schemes too (netstead convert s3://… ./local.parquet). - Architecture — Package → Engine → format dispatch.
- API reference —
Package.write,Package.from_source.