Zingg Python API For Entity Resolution Just Hit the Top 10% of PyPI

Engineering
August 4, 2026

I checked our ClickPy dashboard this week the way I check most metrics these days. Half out of habit, half hoping for a nice surprise. This time I got one.

Zingg now ranks #101,500 overall on PyPI, which puts us in the top 10% of every Python package that exists. Not top 10% of entity resolution tools. Not top 10% of data engineering libraries. Top 10% of all ~1 million packages sitting on the Python Package Index, from requests down to the one-off script published in 2019 and forgotten.

We are super pumped to earn that silver badge on ClickPy’s own ranking tier. And I want to spend a minute on why this particular number is worth a blog post, because “we moved up a leaderboard” is usually the most boring sentence a founder can write.

Why a ranking number, and not just a download count

Raw download counts are easy to inflate — CI pipelines, mirrors, bots re-pulling the same package thousands of times a day. A percentile ranking against the entire PyPI catalog is harder to game, because it’s relative to everything else people are actually choosing to pip install. Sitting in the top decile means Zingg’s batch-based entity resolution engine isn’t a niche tool three teams use internally — it’s being picked up, repeatedly, ahead of the overwhelming majority of what’s published.

Here’s what that install activity actually looks like:

77,000 downloads is the kind of number that only shows up when a tool gets adopted, sits in someone’s requirements.txt, gets pulled into CI, gets reinstalled on a new cluster, gets recommended in a Slack channel by someone who isn’t us. Nobody accidentally installs the same open source entity resolution library 77,000 times.

The GitHub side tells the same story

Downloads are one signal. What happens around the code is another, and it’s arguably the more honest one — because showing up in an issue thread or opening a pull request takes actual effort.

591 pull requests and 573 issues aren’t vanity metrics — they’re a maintenance and trust signal. Every issue is someone hitting an edge case in a real pipeline and caring enough to report it instead of quietly switching tools. Every one of the 31 distinct PR contributors is someone who read the codebase closely enough to change it. That’s a very different kind of validation than a star count, and it’s the kind that tends to compound: an active issue tracker and a real PR history are exactly what a data engineer checks before betting a production MDM pipeline on an open source dependency.

What this actually reflects

We built Zingg’s core matching engine to run natively on the platforms teams already use - Spark, Databricks, Snowflake, Fabric, AWS Glue, GCP. We have kept Zingg architecturally pure to ensure existing data pipelines can run with entity resolution as a first class component, not a bolt on separate source of truth. Entity resolution shouldn’t require ripping out your existing data stack to get deterministic and probabilistic matching that works at scale. Zingg's adoption shows that our approach is correct.

But thats just one thing. I believe there is much more to Zingg than being warehouse native entity resolution. The core strength of Zingg is its simplicity for the end user. Zingg completely abstracts the complex blocking and matching functionality from the end user. Why should a data team worry about field level weights and algorithms, match rules, hyperparameters and other stuff that is needed to build, run and scale entity resolution? Why whould they worry about Spark and Snowpark settings? A 10-15 line program on the Zingg Entity Resolution Python API can achieve all that with ease. Thats by design, and something we are really proud of.

What’s next

We’re extending that same identity-resolution core into a streaming layer for teams that need identity resolved inline, at event time — onboarding flows, order systems, patient-event pipelines, without touching the batch architecture that’s already earned this track record.
‍
We are also actively working on a pure python package for Zingg, without any JVM dependencies. More on that soon.

For now: thank you to everyone who’s pip installed, filed an issue, opened a PR, or just quietly kept Zingg in a pipeline that’s still running. Top 10% of a million packages doesn’t happen without you.

- Sonal

Recent posts