Your data doesn’t live in the “right” format. Neither should your entity resolution

Engineering
August 11, 2026

Here’s a message we get a lot, in one form or another: “We love the idea of Zingg, but half our customer list is still sitting in an Excel sheet someone’s been maintaining since 2019. Do we need to convert everything to CSV or Parquet before we can even try you?”

The honest answer is no. And the reason why is one of our favorite things about how we built Zingg.

Zingg was never really a “CSV tool”

When we designed Zingg’s input and output pipes, we made a deliberate choice: build on Spark’s own reader/writer abstraction instead of writing our own parsers for every file type under the sun. Spark already knows how to talk to CSV, Parquet, Avro, JDBC, Delta and JSON, and through its data source ecosystem, dozens of formats we never had to think about ourselves.

That decision quietly pays off every time someone shows up with a format we “don’t support.” Because more often than not, we already do. We just haven’t said so.

Case in point: Excel

A few weeks ago, a contributor to the Zingg open source project picked up an old open issue asking us to test Zingg against .xls/.xlsx files. What he found (and what we love) is that there was nothing to build. No new pipe class. No parser. No core code touched at all.

He simply pointed Zingg’s config at the com.crealytics.spark.excel format, a Spark data source for Excel files, and it just worked. Zingg read the example febrl test dataset straight out of an .xlsx file, ran the full match pipeline against it and wrote the clustered, deduplicated output back out to .xlsx. There were unicode characters, but that worked too. Sixty-five rows in, sixty-five rows out, schema intact, nothing lost in translation.

That’s the whole story. Add a jar. Point a config at a format string. Done.

Why this matters more than it sounds like it should

If you work in MDM or entity resolution, you know that “real” data rarely arrives in the clean, pipeline-friendly shape we’d like. Vendor lists live in spreadsheets. Sales ops exports customer rosters to Excel because that’s what the CRM spits out. Someone’s regional office still tracks records in a workbook nobody wants to touch.

The traditional answer has been: convert it first, then resolve it. Stand up an ETL step, normalize to CSV or a warehouse table, and only then let your entity resolution engine near it. Every one of those steps is a chance to lose fidelity, introduce a scheduling dependency, or just add a Tuesday you didn’t need.

Because Zingg’s pipes are Spark-native, that conversion step is optional, not mandatory. If Spark can read it, Zingg can resolve entities against it, directly, without us writing a single line of format-specific code. Excel today. Something else tomorrow. The pattern holds because the architecture was built to let it hold.

Try it yourself

The contribution is up as updates to the issue itself. It includes the febrl .xlsx test file, an example config for both reading and writing Excel, and documentation on adding the spark-excel jar to your setup.

If you’ve got messy spreadsheets full of duplicate customers, vendors, or records you’ve been meaning to clean up, just point Zingg at it.

And if there’s a format you’re stuck with that we haven’t talked about yet, open an issue. There’s a good chance the pipe already exists. We just haven’t tried it yet.

Recent posts