Improved

Company Object Accuracy - Interim Guidance

Update, September 8, 2026: corrected data is live for 25 delivery locations. First wave delivered September 2 (21 buckets), second wave September 7 (glassdoor, SEO / website traffic, website intelligence, work milestones / job change). One remains. See Status of corrections below for exactly which buckets are done and what (if anything) you need to do.

We identified an accuracy issue affecting the company object attached to delivered signals. For signal types not yet corrected (listed below), the enriched company attributes we layer onto signals may still be unreliable.

Status of corrections (as of September 8, 2026)

How corrections were delivered. Corrected files were written in place: same bucket, same run folder, same file names, same signal IDs. We did not create new run folders or new file versions. Only the company / contact object changed; the signal event data and your dedup keys are untouched. This means:

  • If you re-read a run folder you already ingested, you get the corrected records. There is no new path to discover.
  • If your loader skips folders it has already seen, you will keep the old company object until you re-pull. For the September 2 buckets, re-pull run folders dated June 15, 2026 or later. For the September 7 buckets (glassdoor, SEO / website traffic, website intelligence, work milestones) every run folder in the bucket was re-enriched, so re-pull the whole bucket.
  • Customers who read from our S3 mirror: the mirror was refreshed the same evening each time (September 2 ~22:00 UTC, September 7 ~19:30 UTC) and matches GCS.
  • Customers we push to directly: we will coordinate the re-push with you on your existing thread.

What changed in the corrected files. Company name, domain, LinkedIn URL, and description were corrected. Employee counts were cleared (set to null) across the corrected files; a verified value will be restored in a follow-up. Industry was left unchanged in the September 2 wave; in the September 7 wave it was re-derived, and for glassdoor and SEO / website traffic that means fewer rows carry an industry than before (the old value was dropped where it could not be verified). Work milestones / job change: the data.company_name field inside the signal was also updated to match the corrected company object.

StatusDelivery location
Corrected, live since September 2gs://autobound-10k-v1, -10q-v1, -8k, -6k-v2, -20f-v2, -financials, -news-v4, -hiring-trends, -hiring-velocity-v1, -github-v1, -patents-v2, -product-reviews-v2, -reddit-company-v2, -twitter-company-posts-v2, -youtube-company-v2, -podcast-company, -podcast-contact, -conference, -producthunt, -federal-contract-award, -new-business-formation
Corrected, live since September 7 (all run folders re-enriched, re-pull the whole bucket)gs://autobound-glassdoor-company-v2, -seo-traffic-v2, -website-intelligence-v1, -work-milestones-v2 (work milestones / job change)
Still to come (keep using the interim workaround)gs://autobound-linkedin-post-company-v2
No correction neededgs://autobound-sec-form-d-funding: the company object on Form D is the filer as named in the filing and was not sourced from the affected table. Verify identity before use still applies (filers are often pooled investment funds without a website).

Measured improvement. September 2 wave, 1,200-signal stratified sample audited field by field against live sources: company LinkedIn URL accuracy 81% → 98%, description 76% → 96%, domain 80% → 91%. September 7 wave, every corrected file compared to its prior version (75M rows) with live LinkedIn page checks on sampled URLs: LinkedIn URL accuracy SEO / website traffic 80% → 90%, glassdoor 92% → 94%, website intelligence 92% → 96%, work milestones 79% → 86%; description fill work milestones 44% → 94%, glassdoor 77% → 96%. We will update this table when the last bucket lands.

Does this affect you? (read this first!)

The signal event and the signal's source identifier are fine. What may be wrong is the enrichment we layered on top of that identifier.

You only need to act if you consume our enriched company attributes. Specifically:

  • Domain-keyed signal types (SEC filings, news, hiring, GitHub, Glassdoor, patents, product reviews, Reddit, Twitter, YouTube, SEO / website traffic, website intelligence, reports & fundamentals): the source company domain on the signal is trustworthy. You are affected only if you use our enriched company name, company LinkedIn URL, description, industry, or employee count instead of enriching from the domain yourself.
  • Company LinkedIn post signals: the reverse. The source company LinkedIn URL is trustworthy. You are affected only if you use our enriched company domain (or the other enriched attributes) instead of enriching from the LinkedIn URL yourself.
  • If you only match signals to your accounts on the source key (domain, or LinkedIn URL for LinkedIn posts) and run your own enrichment - this issue does not touch your workflow at all.

Affected window, by signal type

The issue began when each pipeline started reading the affected enrichment table. Deliveries before the date below are not affected for that type.

Affected sinceSignal typesDelivery location
June 2, 2026work milestones / job change (company object only)gs://autobound-work-milestones-v2
June 15, 2026SEC 10-K, 10-Q, 8-K, 6-K, 20-Fgs://autobound-10k-v1, -10q-v1, -8k, -6k-v2, -20f-v2
June 15, 2026news (all versions)gs://autobound-news-v2, -news-v3, -news-v4
June 15, 2026earnings transcripts, reports & fundamentalsgs://autobound-earnings-transcripts-v2, -financials
June 15, 2026hiring trends, hiring velocitygs://autobound-hiring-trends, -hiring-velocity-v1, -signal-delivery
June 15, 2026GitHub, Glassdoor, patents, product reviews, Reddit, Twitter (company + contact posts), SEO / website traffic, website intelligence, company LinkedIn posts, contact LinkedIn postsgs://autobound-github-v1, -glassdoor-company-v2, -patents-v2, -product-reviews-v2, -reddit-company-v2, -twitter-company-posts-v2, -twitter-contact-posts-v2, -seo-traffic-v2, -website-intelligence-v1, -linkedin-post-company-v2, -linkedin-post-contact-v3
July 13, 2026LinkedIn commentsgs://autobound-linkedin-comments-contact-v2
July 15, 2026YouTube (company)gs://autobound-youtube-company-v2
August 25 / 28, 2026conference speaker, SEC Form Ddays-old window; verify identity before use
Not affectedweb intent (no company object), YouTube contact-level, employee growth-

Employee growth is confirmed unaffected two ways: its pipeline has no code path into the affected table, and its delivered company attributes come from its own LinkedIn snapshots (verified by comparing delivered records against the table - they do not match).

Which fields to trust, on the object itself

The company object (full schema here) is identical across signal types - what changes is which field is the trustworthy source key.

Example 1 - a domain-keyed signal (SEC filings, news, hiring, GitHub, Glassdoor, patents, product reviews, Reddit, Twitter, YouTube, SEO / website traffic, website intelligence, reports & fundamentals):

"company": {
  "domain": "acme-robotics.com",          // TRUSTED - the source key this signal resolved to. Match on this.
  "name": "Acme Robotics",                 // enriched - may be wrong until corrected
  "linkedin_url": "https://www.linkedin.com/company/acme-robotics",  // enriched - may be wrong or a subsidiary/regional page
  "industries": ["Industrial Automation"],  // enriched - may be wrong
  "employee_count_low": 51,                 // enriched - may be wrong
  "employee_count_high": 200,               // enriched - may be wrong
  "description": "Acme Robotics builds..."  // enriched - may describe a DIFFERENT company
}

To be safe today: keep matching on domain, and re-derive everything below it from that domain (your own enrichment or our enrichCompany endpoint).

Example 2 - a company LinkedIn post signal (the keys reverse):

"company": {
  "linkedin_url": "https://www.linkedin.com/company/acme-robotics",  // TRUSTED - the page this post was collected from. Match on this.
  "domain": "acme-robotics.com",            // enriched - may be wrong until corrected
  "name": "Acme Robotics",                  // enriched - may be wrong
  "industries": ["Industrial Automation"],  // enriched - may be wrong
  "employee_count_low": 51,                 // enriched - may be wrong
  "employee_count_high": 200,               // enriched - may be wrong
  "description": "Acme Robotics builds..."  // enriched - may describe a DIFFERENT company
}

Same object, opposite trust: the post was gathered from that LinkedIn page, so linkedin_url is ground truth and domain is the derived field to re-check.

Interim workaround: re-enrich from the trusted key

For the signal types still to come in the status table above: each affected signal already carries a reliable identifier. Key off it and do not rely on the delivered enriched fields until that bucket is marked corrected:

Re-enrich usingSignal types
Company domain (the source domain on the signal)SEC filings incl. 20-F, news, reports & fundamentals, GitHub, hiring trends, patents, podcasts, product reviews, SEO / website traffic, Glassdoor, website intelligence, Reddit, Twitter, YouTube
Company LinkedIn URLCompany LinkedIn posts
Verify identity before useSEC Form D, conference speaker

What happens next

  1. Rebuild (complete September 1). The enrichment merge was rewritten so that company identity (name, domain, LinkedIn URL) is verified first - across sources and against the company's live LinkedIn page - and descriptive attributes are only accepted from sources that agree with that identity. Fields that cannot be verified ship blank rather than wrong, with the suppression logged.
  2. Validation before anything ships. Every rebuild passes three gates: an automated completeness check (no value dropped without a logged reason), a bulk text-consistency sweep across a 20,000-row sample, and a Groundtruth audit of a stratified 100-company sample against live sources, which must land within 3 points of the measured target bands (name ~87%, description ~98%, LinkedIn URL ~96%, industries ~96%, employee counts ~97%).
  3. Live verification layer. For every company appearing in delivered signals, we verify the LinkedIn URL against the live page and take employee counts and industry from that verified page rather than from any vendor file.
  4. Corrections upserted into your delivered data (first wave delivered September 2). Every corrected record keeps its original signal ID; only the company/contact objects change - your dedup keys and event data are untouched. Files are corrected in place at their existing paths (see the status table above); a correction manifest is available on request for teams that prefer to re-ingest deltas explicitly. Remaining buckets follow as each pipeline is re-run.
  5. Before/after accuracy published for review - per field, per signal type. Headline numbers are in the status section above; the full per-type breakdown is available on request.
  6. Continuous auditing. Groundtruth samples every future delivery, so regressions of this class are caught before data ships.

A detailed engineering retrospective is available on request.

What happened, briefly

On August 28 an internal quality review of the job-change signal surfaced company records whose attributes didn't belong to the company on the signal.

The trigger was a new upstream data file we had recently begun ingesting. The file is genuinely excellent at what it was brought in for - contact-level detail, with remarkably accurate job titles, current company names, and detailed work histories. The oversight was internal, not the vendor's: we assumed the file's company-domain-to-LinkedIn-URL mapping carried the same accuracy as its contact fields, and that specific slice was never independently validated before the file was upserted into our existing company enrichment table. It turned out to be the file's weakest link - and our merge logic made a weak link expensive. The merge trusted each source's claimed company identifiers without verifying, across sources or against live pages, that they described the same company, and it selected each output field independently by source priority. A mislabeled company binding could therefore contribute fields to another company's record, blending attributes across entities into records that looked plausible individually. Once the enrichment table began filling the company object on shipped signals in mid-June, the pollution shipped with them.

Groundtruth - our internal audit engine that checks each field of a delivered signal against live authoritative sources (the company's own website, LinkedIn, SEC filings) - had been in our engineering pipeline for months and launched a few weeks before this incident; its integration into the delivery flow was still in progress, so it was not yet gating every drop. Once the issue was detected we turned Groundtruth on the problem directly: within 48 hours we had audited samples across every delivery, quantified accuracy per field per upstream source - including the validation slice that was originally skipped - redesigned the merge (verified identity first; attributes only from sources that agree with that identity; a blank over a wrong value), and measured the new logic at 87–98% field accuracy on live samples. Completing Groundtruth's integration so it gates every delivery is now the top of that roadmap, alongside a standing rule this incident bought: no source field enters the enrichment table until that specific field's accuracy has been measured - reputation, even earned reputation, doesn't transfer between fields.

We communicated this guidance to customers on August 31, delivered corrections on September 2 and September 7, and will keep this entry updated until the last bucket is corrected. Last updated September 8, 2026.