How to combine Elite Prospects, Hudl and Sportlogiq data into one place
Why hockey data sources never line up, and the four approaches clubs actually use to get them into a single analysis-ready table.
The short answer
There is no shared id between hockey data providers, so combining them means building an entity-resolution layer that matches players, teams and games across sources. Clubs do this one of four ways: manual spreadsheet reconciliation, a scripted pipeline maintained in-house, a general-purpose ETL tool with custom connectors, or a hockey-specific integration platform. The right choice depends on how many sources you have and whether you can afford an analyst spending their week on maintenance.
A club buys a league data subscription, a video platform, a tracking provider and an athlete management system. Each is a good product. Then someone asks a question that needs two of them at once — *how do our zone exits look for the players who were flagged as high load last week?* — and the whole thing falls apart.
This post is about why that happens and what the actual options are.
Why the sources do not line up
The problem is almost never the quality of the data. It is that nothing connects it.
There is no shared player id
Elite Prospects has its own player ids. Hudl has its own. Sportlogiq has its own. The NHL API has its own. None of them reference each other, because each system was built before any standard existed, and those identifiers are load-bearing inside each product.
So joining them means matching on what is left: names, birthdates, teams and seasons. Names are where this gets unpleasant.
- The same player appears as Mikael Backlund, Backlund, Mikael, M. Backlund and Mikael BACKLUND
- Transliterated names vary by source — Kaprizov and Kaprisov are the same player
- Diacritics get stripped inconsistently, so Jiří becomes Jiri in one feed and Jir in another
- Two players genuinely share a name and you need birthdate or team to tell them apart
Games and teams have the same problem
A single game has a different identifier in every system. Team names drift too — a source might carry Frölunda HC, Frolunda, Frölunda Indians and an abbreviation, sometimes within the same export, depending on which season the row belongs to.
Definitions quietly disagree
This is the one that produces wrong answers rather than obvious errors. Two sources both have a column called shots, but one counts attempts toward the net and the other counts shots on goal. Both are correct by their own definition. Joined without checking, they produce a number that is confidently wrong — and nothing in the pipeline will warn you.
The four approaches
1. Manual spreadsheet reconciliation
Export CSVs, paste them into a workbook, match names with lookups, fix the mismatches by hand.
This genuinely works for a one-off question about one team. It is the right call for a question you will ask once. It stops working the moment the answer needs refreshing weekly, because the manual matching is repeated in full every single time, and any fix someone made last week is gone.
2. A scripted pipeline, maintained in-house
An analyst who can code writes Python that pulls each source, normalises names, and writes to a database. This is a real solution and plenty of clubs run on it.
The catch is not the build — it is that the pipeline is never finished. Providers change schemas without notice, rotate credentials, alter export formats and occasionally retire endpoints. Each break is small. Together they turn into a standing part-time job, and it lands on the person you hired to analyse hockey. When that person leaves, the club inherits an undocumented pipeline nobody else understands.
3. A general-purpose ETL tool
General-purpose ETL tools solve exactly this shape of problem — and solve it well — for marketing and sales data. The difficulty is that none of them ship hockey connectors, so you are writing custom sources anyway, and you are now maintaining those custom connectors *inside* someone else's framework.
They also do not help with the hockey-specific part. Entity resolution across hockey providers is domain knowledge, not a generic transform.
4. A hockey-specific integration platform
A platform that ships maintained connectors for hockey sources and handles identity resolution as a built-in concern rather than something each club reinvents. Connector maintenance moves off club staff, and the matching logic improves for everyone as it gets corrected.
This is the category The Hockey Brain is building in, so treat that as a disclosed interest rather than a neutral assessment. It is worth naming what you give up: a dependency on someone else's roadmap, and less control than a pipeline you wrote yourself.
How to choose
A reasonable rule of thumb:
- One or two sources, occasional questions — spreadsheets are fine. Do not over-engineer this.
- Two or three stable sources, someone technical who owns it — an in-house pipeline is usually the right call.
- Four or more sources, or nobody who wants to own maintenance — the maintenance burden is the deciding factor, not the build effort.
The honest signal that you have outgrown your current approach is not that the data is wrong. It is that someone capable is spending a meaningful share of their week on reconciliation, and questions stop being asked because the answer takes three days to assemble.
If you are building it yourself
A few things worth knowing before you start:
- Build the identity layer first. A table mapping each source's player id to one internal id is the foundation. Get it wrong and every downstream number inherits the error.
- Store the raw extract before transforming. When a number looks wrong six months from now, you need to see exactly what the provider sent.
- Keep a manual override table. Automated matching will reach roughly 95% on names. The remaining 5% needs a human decision, and that decision must survive the next run.
- Write down your definitions. When two sources disagree about what a shot is, the resolution has to live somewhere other than in one person's memory.
- Check your licences. Most provider agreements govern how data may be stored and redistributed. Worth reading before it is in a warehouse.
Frequently asked questions
Can I export data from Elite Prospects and Hudl into one spreadsheet?
Yes, both allow exports, but the exports will not join cleanly. Player names are formatted differently, there is no shared player id, and the two systems disagree about game identifiers. You can reconcile a single team's roster by hand in an afternoon; doing it for a full league every week is where it stops being viable.
Why do hockey data sources not share player ids?
Each provider built its own identifier system before any standard existed, and those ids are core to their product. There is no league-wide or sport-wide identity standard in hockey the way there is in some other sports, so every integration has to resolve identity itself using names, birthdates, teams and seasons.
What is entity resolution in hockey data?
Entity resolution is the process of deciding that a player, team or game in one data source is the same real-world entity as a record in another source. For hockey it typically combines name normalisation, birthdate matching, team and season context, and a manual review queue for ambiguous cases such as players with common names or transliterated names.
Should a club build its own hockey data pipeline?
It is reasonable when the club has two or three stable sources and someone technical who owns it. It becomes expensive once there are five or more sources, because providers change schemas, rotate credentials and alter export formats on their own schedule. The build is rarely the problem; the ongoing maintenance is.
Tired of reconciling exports?
We are building maintained connectors for hockey data sources. Early clubs decide what ships first.