Campaign finance data should be public by law. But being legally public and actually being understandable are two very different things.
In North Carolina, decades of campaign contribution records were technically available, sitting on the NC State Board of Elections (NCSBE) website, waiting for anyone who cared to look. The problem wasn’t access. It was chaos: 50 million transactions, two incompatible data formats, and thousands of donors and recipients recorded under dozens of slightly different names. “New Belgium Brewing Co.” and “New Belgium Brewery PAC” were the same entity, but the data didn’t know that. Neither did any analysis built on top of it.
The result: nobody could reliably answer the most basic question in political accountability, who is spending money on what, and how much?
This is the story of how CrossroadsCX, working with the North Carolina Free Enterprise Foundation, used Zingg to solve that problem — and what it took to make 50 million transactions actually tell the truth.
The NCSBE had campaign finance data going back years, but it existed in two fundamentally incompatible forms.
The first was structured digital records, machine-readable files that could be queried directly. The second was a large archive of scanned PDF reports from campaigns that had submitted paper filings instead of digital ones. Those PDFs had to be manually transcribed before they could be analyzed at all.
Even after transcription, the bigger problem surfaced: the data was riddled with manual entry errors, inconsistent formatting, and no standardization across submissions. Different people filing on behalf of the same organization or even the same person filing at different times, would spell names differently, abbreviate differently, use different addresses, or include or drop middle initials.
This is where previous transparency efforts had stalled. As Jimmy Steinmetz, co-founder at CrossroadsCX, put it plainly:
“Entity resolution was just too much of a bear to successfully pull off a project like this.”
The problem wasn’t will. It was that matching entities across messy, inconsistent, multi-source data at scale is genuinely hard and getting it wrong doesn’t just produce inaccurate dashboards. It produces misleading ones. If “Robert Johnson” and “Bob Johnson, Esq.” and “R. Johnson” are treated as three separate donors, you fundamentally misread the funding landscape.
Before understanding the solution, it helps to appreciate the scale.
The full NC campaign finance database contains roughly 50 million transactions. Not all of those involve named entities that need resolution, many are straightforward line items. But after filtering to the records that required entity matching, the team was working with approximately 500,000 unique organizations, contributors, recipients, PACs, companies, and individuals, each of which might appear in any number of different forms across the dataset.
That’s not a task you can solve with a spreadsheet lookup or a simple string-matching rule. You need something that understands when two differently-spelled names are probably the same entity, and can make that judgment across hundreds of thousands of records simultaneously.
The scope of the entity variation problem becomes vivid with a single example.
Zingg identified 22 distinct ways that “Facebook” appeared as an expense recipient in the North Carolina campaign finance data. Facebook as a vendor for political ad spending is common enough that it appears frequently but submitted by many different campaigns, with many different data entry habits, under entries like “Facebook Inc.”, “Facebook, Inc.”, “Facebook Ads”, “Meta (Facebook)”, “FB Advertising”, and so on.
Without resolution, those 22 entries look like 22 different vendors. With resolution, they collapse into one: Facebook. And the true scale of ad spending through that channel becomes visible for the first time.
This dynamic plays out across thousands of entities in the dataset, not just tech companies, but local businesses, law firms, consulting agencies, and individual donors whose names appear in slightly different form each time they contribute.
The CrossroadsCX team deployed Zingg on Snowflake to handle the entity matching. What made Zingg the right tool for this problem was its ability to match across multiple fields simultaneously, not just name, but also address, city, and other available metadata.
As Steinmetz explained:
“We’re not just doing this based on name. We’re also matching based on address, city, and other fields.”
This matters enormously in practice. Names alone produce both false positives (two different “John Smith” donors who happen to share a name) and false negatives (the same donor whose name was abbreviated differently in different filings). Combining name with address and other signals dramatically improves accuracy.
Zingg’s ML-based approach learns the right matching weights from human-labeled examples rather than requiring a team to manually write and maintain matching rules. For a dataset with this many variations, a rules-based approach would require constant maintenance as new variations appeared. Zingg’s model generalizes.
The output of Zingg’s matching is clusters, groups of records that all refer to the same real-world entity, along with a standardized, cleaned canonical name for each cluster. When “New Belgium Brewing Co.” and “New Belgium Brewery PAC” both get resolved to the master entity “New Belgium Brewing Company,” every transaction linked to either variant is now attributable to the same entity.
Across the full dataset, Zingg identified approximately 17,000 such clusters, 17,000 cases where what looked like different entities were actually the same one. That’s 17,000 places where the data had been silently misleading anyone who tried to analyze it.
The full pipeline CrossroadsCX built combined several cloud-native tools into an end-to-end data platform.
Data ingestion started with Node.js utilities that pulled digital transaction records directly from the NCSBE website, while scanned PDF reports were transcribed and converted to machine-readable format. All raw files landed in Google Cloud Storage.
Processing ran through Google Cloud Functions, triggered by file uploads or scheduled intervals. This is where the three-stage cleansing happened:
After cleansing and resolution, data flowed into Snowflake via Snowpipe for continuous automated ingestion. Snowalert (running on Google Kubernetes) monitored the pipeline and flagged failures.
A key architectural detail: the team used transaction IDs as a bridge between the resolved entities and the raw transaction records. Once Zingg selected a master entity name (say, "New Belgium Brewing Company"), that canonical name could be mapped back to every individual transaction, regardless of how the entity originally appeared, ensuring that all $X attributed to that entity was correctly aggregated.
The analytic layer sat entirely in Snowflake and consumed Zingg’s output directly. Tableau Public and d3.js powered the public-facing dashboards, enabling anyone journalist, researcher, or curious citizen to explore spending patterns interactively.
The numbers tell the story clearly.
Zingg’s entity resolution processed 500,000 unique organizations out of a dataset of 50 million transactions, identifying ~17,000 clusters of records that referred to the same entity under different names. The result: $250 million in campaign financial transactions that had previously been fragmented, double-counted, or simply unattributable now resolved into a coherent, accurate picture of North Carolina political spending.
The public dashboards built on this data let users do things that were previously impossible:
For a project explicitly aimed at democratic transparency, that last point matters most. Campaign finance analysis is only as good as the entity data underneath it. If “Friends of Candidate X” and “Friends for Candidate X Committee” are treated as separate entities, the true picture of that candidate’s funding is invisible.
The NC campaign finance project is a compelling proof of concept for something broader: that entity resolution isn’t just a commercial data problem.
Civic datasets, voter rolls, property records, court filings, government contractor databases, share the same structural characteristics as commercial customer data. They’re compiled from multiple sources, submitted by humans, inconsistently formatted, and riddled with the natural variation that comes from having no single canonical system of record. They also have enormous stakes: inaccuracies don’t just affect a marketing campaign’s conversion rate, they affect accountability, transparency, and democratic oversight.
The pattern that made this project work, combining a multi-field ML matching engine with a cloud-native pipeline and a clear canonical record strategy, is reproducible. Any dataset that involves named entities, submitted by different parties over time, with no enforced standardization, is a candidate for the same approach.
For CrossroadsCX and the North Carolina Free Enterprise Foundation, that meant turning 50 million disorganized transaction records into a usable transparency tool. For others, the same playbook might apply to public health records, campaign donor registries, or government procurement databases.
The data existed. The challenge was always making it tell the truth.
Zingg is open source and available on GitHub. For enterprise deployments on Snowflake, Databricks, or other cloud platforms, reach out to the Zingg team.
If you’re working on a public data transparency project, civic, nonprofit, or otherwise, we’d love to hear from you.
Join the Zingg community on Slack and share what you’re building.
Case study details sourced from CrossroadsCX’s public documentation and their work with the North Carolina Free Enterprise Foundation. All technical architecture credits to Jimmy Steinmetz and the CrossroadsCX team.