Quickstart — corral ¶
When to use this ¶
You want to load + validate + (optionally) spatially scope a Frictionless data package — any spec, not just GMNS — and see what corral gives you in five minutes. If you’re working with GMNS networks specifically, the netstead quickstart is the better starting point; everything here applies underneath that surface.
Quick example ¶
The fastest way to confirm corral is wired up is to load a real Frictionless package, run the four-pass validator, and print a one-line summary. The example below uses the bundled Leavenworth GMNS fixture as a stand-in for “any Frictionless package” — the API is identical regardless of spec.
Then in Python:
from corral import Package
from netstead.fixtures import leavenworth # any Frictionless dir works here
pkg = Package.from_source(leavenworth.csv_dir(), spec=leavenworth.spec_path())
report = pkg.validate()
print(f"{len(pkg.tables)} tables, {len(report.issues)} validation issues")
Expected:
You just loaded a Frictionless data package through the default ibis + DuckDB engine, then ran the structural / schema / FK / sync-state validator in a single call. Nothing was eagerly materialised — pkg.tables["link"] is still a lazy expression.
Step-by-step ¶
1. Install ¶
The default install ships everything for the four-pass validator below. Pick your tool:
Optional extras (polars / pandas / s3 / gcs / azure / keyring / mcp) opt in to alternative engines, cloud backends, and the AI surface — see the install guide for the full list.
2. Load a Frictionless package ¶
Package.from_source accepts a local path, an s3:// / https:// / duckdb:// URL, or a directory of CSV / Parquet / DuckDB / zip-CSV files. The example below loads the bundled Leavenworth fixture, which is a CSV directory. Because a CSV-only directory doesn’t carry a datapackage.json manifest, pass spec= explicitly to tell corral which schema to validate against.
from corral import Package
from netstead.fixtures import leavenworth
pkg = Package.from_source(
leavenworth.csv_dir(),
spec=leavenworth.spec_path(),
)
You should see something like:
Substitute any other Frictionless package — your own spec, a GTFS feed converted to a Frictionless package, or a cloud-hosted parquet partition — and the rest of the steps below are identical.
3. Validate and read the report ¶
Validation runs four passes in one call — structural (required resources present), schema (field types and constraints), foreign-key (cross-table integrity), and sync-state (have FKs gone stale since a previous edit?). The result is one ValidationReport with severity-graded issues.
report = pkg.validate()
print(f"{len(report.issues)} issues across {len(pkg.tables)} tables")
for issue in report.issues[:5]:
print(f" {issue.severity.value:9} {issue.code:30} {issue.message[:60]}")
In a Jupyter notebook, evaluating report on its own renders an HTML card grouped by severity. In a script, you can also serialise to JSON for CI:
import tempfile, pathlib
with tempfile.TemporaryDirectory() as d:
report.to_json(pathlib.Path(d) / "validation.json")
4. Scope to a spatial subset ¶
If your package has a geometry column (any WKT or WKB), corral.dataset.view.from_bbox returns a lazy view filtered to features inside a bounding box. The example below scopes the Leavenworth geometry table (which carries the inline WKT) to a small bbox around the historic core.
from corral.dataset.view import from_bbox
geometry = pkg.tables["geometry"]
scoped_geom = from_bbox(geometry, -120.665, 47.594, -120.655, 47.600)
print(f"scoped: {scoped_geom.count()} geometries in bbox")
The filter pushes down to DuckDB — only the matching rows are read.
5. Write the package out ¶
pkg.write(dest) round-trips the package to a new location. The output format is inferred from the extension: .parquet for a Parquet directory, .duckdb for a single-file DuckDB, or a directory for CSV.
import tempfile, pathlib
with tempfile.TemporaryDirectory() as d:
pkg.write(pathlib.Path(d) / "out.parquet")
The datapackage.json manifest is regenerated alongside the data, so the written-out package is itself a valid Frictionless package.
Common variations ¶
Each accordion below is one alternative to the defaults used above. The first (most common) is expanded; the rest are collapsed — open the one that matches your situation.
Switch the backend engine to pandas or polars
Default is IbisEngine (lazy, DuckDB-backed). Switch to pandas for eager DataFrame ergonomics, or polars when you want fast in-memory analytics.
Read from S3 with credentials
Credentials resolve via a cascade: keyword arg → CORRAL_CRED_<host>_TOKEN env var → OS keyring → .netrc. No code changes needed if the environment is set up.
Load only a subset of tables
Pass tables= to materialise only the tables you need — useful when the package is large and you only care about a few resources.
Pass the spec from a URL or another package
spec= accepts a local path, a URL, or an already-loaded Spec object. Useful for sharing a spec across many sibling datasets.
Pitfalls ¶
- No
datapackage.jsonin your directory → passspec=explicitly. CSV-only fixtures and ad-hoc directories don’t carry a manifest, socorralcan’t infer the schema. Pointspec=at the canonicaldatapackage.jsonfor your data. See Frictionless data packages for the spec-vs-data distinction. - Credential cascade order matters. Per-call
credentials=always wins; otherwiseCORRAL_CRED_<host>_TOKENenv var, then OS keyring, then.netrc. If you’ve stashed a token in two places, the higher-priority one is the one that gets used. - Lazy by default. Tables are lazy ibis expressions over DuckDB — nothing materialises until you call
.to_pandas()/.to_polars()/.count(). Push predicates down and only pull the rows you need into Python.
See also ¶
- Cookbook — task-oriented recipes (read from S3, convert formats, spatial scope).
- API reference — every public symbol with stable anchors.
- Architecture — defaults, rationales, extension points.
- Frictionless data packages — the spec
corralbuilds on.