Shovel Research

We tested 89,729 GLEIF corporate relationships

GLEIF Level 2 data is useful for finding possible accounting-consolidation links. An active, published row is not enough to accept a current parent.

August 27, 2026 · by Shovel
Abstract company records connected in a relationship graph, with one relationship accepted, one rejected, and many unresolved
One relationship in our 60-row test cleared the automatic-use bar, one failed it, and 58 did not contain enough evidence to decide.

GLEIF, the Global Legal Entity Identifier Foundation, looks at first like the answer to a problem every company-data team eventually meets. It gives legal entities stable identifiers, publishes relationship files, records dates and validation fields, and makes the data available for reuse. If you need to connect company names across government records, the prospect of a global corporate graph is hard to ignore.

We nearly used it that way. Then we tested it.

Is GLEIF reliable?

Our conclusion is narrower than "GLEIF is unreliable." The files are structurally clean and unusually well documented. GLEIF is useful as a source of proposed, dated accounting-consolidation relationships. But an ACTIVE, PUBLISHED row must not become an accepted current parent relationship automatically.

What GLEIF Level 2 actually states

GLEIF Level 2 answers an accounting question: which legal entity directly or ultimately consolidates another entity's financial accounts. The relationship we evaluated was IS_DIRECTLY_CONSOLIDATED_BY.

That is not a general statement of legal ownership, operational control, brand ownership, franchising, facility ownership, or every member of a corporate family. Treating it as a universal "who owns whom" graph changes the meaning of the source before any quality problem enters the picture.

The useful part is that each record names the child and parent by LEI and can include relationship periods, registration status, validation basis, update dates, and a validation reference. Any system that ingests the data should preserve those qualifications.

How we tested it

We used the GLEIF Golden Copies published at 2026-08-26 16:00 UTC and measured our Standard Record and Safety Incidents identity inventory at 2026-08-26 23:25:27 UTC.

The tested population was all 89,729 direct-consolidation relationships whose relationship status was active and registration status was published.

The evaluation had three parts. We measured structural, evidence, lifecycle, and period quality across all 89,729 rows. We projected exact identifier overlap into our own company inventory using SEC CIKs only, with no name, address, or jurisdiction similarity. Then we drew a deterministic 60-row adversarial sample, ten rows from each of six quality strata, and reviewed whether each row could safely become an automatic current edge.

The sample seed was standardrecord-issue-12-gleif-holdout-v1. The sample was stratified, not random over the full population. It was designed to expose failure modes, not estimate how often each failure occurs.

The files were structurally clean

All 89,729 rows represented unique child-parent pairs. We found no duplicate pairs, self-relationships, children with multiple direct parents in the tested slice, or two-node cycles. General legal entities appeared at both endpoints in 86,530 relationships, or 96.43%. GLEIF reported fully corroborated validation for 50,804 relationships, or 56.62%.

Those are real strengths. Stable LEI endpoints and explicit provenance fields make GLEIF a strong place to find candidate hierarchy statements.

Why active and published was not enough

A future accounting-period end is not necessarily an error. A lapsed LEI does not prove a relationship ended. Both still require interpretation before a product tells a reader that one company is the current parent of another.

Independent verification was also thin. Only 8 of the 60 sampled rows supplied a validation URL. When checked on August 26, three returned HTTP 200, two returned 403, one returned 404, one returned 500, and one timed out.

The 60-row adversarial sample

We selected ten rows from each of these populations:

Eligibility for automatic current use, 60-row adversarial sample
Eligible
1
Rejected
1
Abstained
58
Stratified across six quality conditions. This is a failure-mode test, not an accuracy estimate for the full file.

The one positive was Canada Life Finance (U.K.) Limited to The Canada Life Assurance Company. Linked 2024 Companies House accounts named Canada Life Assurance as the parent of the smallest consolidating group containing the child.

The one negative was Vontobel Asset Management S.A., Niederlassung München to Vontobel Asset Management S.A. The GLEIF row was active and published, but the child endpoint was inactive and retired, and the accounting period ended in 2016.

The other 58 rows were abstentions. The published record did not contain enough current, direct, checkable evidence to accept or reject the relationship under our evidence rules.

Two examples show what an abstention means. The linked accounts for LGIM Maturing Buy & Maintain Credit Fund 2040-2054 and Legal and General Assurance (Pensions Management) Limited placed both in a wider subsidiary population, but did not establish the claimed direct consolidating-parent edge. The Wing Re Inc. relationship to Swiss Re Life & Health America Holding Company was entity supplied and included no validation reference.

This does not mean GLEIF is 1-for-60 accurate. We measured eligibility for automatic current use, not whether a relationship had ever been true, and we deliberately oversampled difficult conditions. The result is that a published row, by itself, usually did not clear our evidence threshold.

It barely connected to the identifiers we already trust

We also measured exact SEC CIK overlap against our live identity inventory. The projection used GLEIF's own SEC registration-authority identifiers. It did not use fuzzy names.

With no relationship resolving both endpoints, GLEIF added no hierarchy edge we could safely connect to our current graph. Using it would first require another endpoint source or a separately measured resolver.

Private-company coverage is less measurable, not automatically better. GLEIF often publishes state or national registration identifiers. Our current inventory holds EINs, CIKs, DOT numbers, and agency facility or establishment IDs, but not a broad state-registration crosswalk. We also do not have a reliable public-versus-private classification, so any claim about private-company coverage would lack a defensible denominator.

How we would use GLEIF

  1. Store each row as a versioned, source-addressable proposed accounting-consolidation assertion, never as an accepted generic parent edge.
  2. Preserve the relationship type, periods, registration lifecycle, validation basis, documents, reference, source timestamp, and both LEIs.
  3. Resolve each endpoint independently through identifier evidence already accepted by the product.
  4. Require current endpoint status and checkable support before automatic promotion. An entity-supplied-only row should abstain unless another source corroborates it.
  5. Calibrate every promotion class against independent labels. For an automatic class, our rule is zero observed false joins and a Wilson 95% lower precision bound of at least 0.99. With no observed error, that requires at least 381 representative decisions.
  6. Treat abstention as a valid result. A coverage target must never weaken the precision gate.

There is no reason for us to prioritize GLEIF ingestion until endpoint coverage improves or a product can use unresolved proposed statements. The next useful investigation is an endpoint-source decision, not another pass at fuzzy name matching.

Reproduce the work

This article consolidates Standard Record employer-matching issue #12 and its handoff. We did not repeat the sampling for this publication. The exact report, sampler, read-only queries, source decision, and organization-ontology ADR are linked below at the commits used for this evaluation.

Get the next finding.

We dig these out of public data. One email when we publish the next one. Nothing else.