The 500-Lead Problem
You searched for "plumbers in Houston" and got 500 results. Names, phone numbers, addresses, websites. A fat spreadsheet of opportunity.
Except some of them are not plumbers.
That is not a guess. In August 2026 we took 12 keyword searches, sampled 360 US businesses, fetched every homepage, and judged each one against the keyword that surfaced it. 72.2 percent were genuinely the business searched for. The rest were an adjacent trade, a different business entirely, or a directory page. Full method and per-keyword numbers are in the study.
The average is the least useful part of that. "Roofing contractor" came back 95.8 percent clean. "Home remodeling contractors" came back 28.0 percent. How precise your keyword is decides how much of your list is real, and nothing warns you which kind you just ran.
"Things Google Maps returned for the search 'staffing agency': two performing arts theatres, a car rental branch, a UPS Store, and a Marine Corps recruiting office."
Why a keyword search returns the wrong trade
Google Maps matches on more than the business category. It matches on text in reviews, on services mentioned in passing, on the name, on proximity to what it thinks you meant. A general contractor whose site mentions roofing surfaces for "roofing contractor". So does a roof cleaning company, a gutter installer, and a supplier who sells to roofers rather than being one.
None of that is a bug. It is what makes Maps useful when you are a consumer looking for someone to fix a leak. It is a problem when you are building a prospect list and every wrong row costs you a real minute of a real salesperson's day.
The obvious fix, and what it costs you
Every Maps result carries a business category, so filter on it. This works better than we expected: filter to the categories the genuine results carried and almost all of the noise disappears. If you are building lists and not filtering on category, start.
But it is not free, and the cost is invisible, which is the worst kind. In our study that same filter also discarded 24 of 205 genuine businesses, 11.7 percent. We read all 24 by hand:
- Big Brand Tire & Service, whose own page reads "Tire Shop & Auto Repair in Phoenix, AZ", is categorised Tire shop. Dropped from an auto repair list.
- J & J Remodeling Team LLC, "a locally owned remodeling company", is categorised Remodeler rather than the Home builder the other genuine results carried.
- A pediatric dentist, an orthodontist and a periodontist, all dropped from a dentist list, because Google files each dental specialism as its own category.
- A plumber categorised Home help. Nobody would have guessed that one in advance.
The category is one self-chosen label from a list of thousands, and it usually describes a specialism rather than a trade. To filter without losses you would have to know the whole family of categories in advance, for every trade you sell into.
What actually works: read the website
The only place a business reliably says what it does is its own homepage. So that is what we check. Every enriched result is read against the keyword that produced it, and marked Yes, No or Unclear, with the sentence from the page that justifies the verdict sitting next to it.
We tested this against the category filter using an independent second rater: a different model, asked a differently worded question, blind to the first verdict, on the same 267 pages. The two raters agreed 89.5 percent of the time.
| Method | Precision | Recall |
|---|---|---|
| Filter on Google category | 83.4% | 87.4% |
| Read the website | 86.4% | 98.8% |
Precision is a tie. Recall is not close. The category filter missed 21 genuine businesses; reading the site missed 2. So the value is not that it catches more junk. It is that it keeps the businesses a category filter silently deletes.
One honest limitation, because it matters: both raters are language models reading the same page text, so they share a method bias we cannot size. A panel of human raters would be the stronger test and we have not run one. Treat the direction as well supported and the exact size of the gap as provisional.
Every verdict quotes the page
A judgement you cannot check is worth nothing, so the quote ships with the verdict or the verdict does not ship. If the model returns a verdict whose supporting quote does not actually appear on the page, we downgrade it to Unclear rather than show it.
That rule earned its keep immediately. In testing, a business whose only web presence was a Facebook page came back as a confident No with the evidence "The page does not provide any information about the company". That is a sentence about the page, not from it. The honest answer there is Unclear: we could not read the site, which is not the same as reading it and rejecting the business.
Blank means the same thing, and never means No. A row we did not evaluate, because it has no search behind it or because the site could not be fetched, stays empty. Every field we ship follows the same three-state rule: a confirmed absence is stored separately from a value nobody checked.
Using it on a real list
Sort or filter the column and work the Yes rows first. On a broad keyword that is the difference between a list that is 28 percent real and a working queue.
Two things worth knowing. Uploaded leads have no search behind them, so there is nothing for the check to judge against; the upload dialog asks what kind of business the file contains so the column is not silently empty. And Unclear is worth a look rather than a delete, because it usually means a thin or JavaScript-only site rather than a wrong business. In our sample 21.1 percent of businesses had a website we could not read at all, which is its own signal if you sell websites.
What this does not do
It tells you whether a business is the kind of business you asked for. It does not tell you whether they can afford you, whether they are in a buying cycle, or whether the owner will like you. That is still your job, and the enrichment fields are there to help you do it: no CRM detected, a slow site, high review count against a low rating, ecommerce with no email marketing.
We used to put a model in front of those fields and have it write a label summarising them. We removed that, because when we finally ran it on real leads, half the sample got the identical label and every piece of evidence was a database column restated back at us. A filter on the column does the same job for free, instantly, and sortably. Relevance is the part that needs a model, because it needs someone to read the prose.
Getting started
Run a search, let enrichment finish, and look at the Relevant column. It costs nothing extra; it is part of enrichment. If you are on a broad keyword, compare the Yes count to your total before you plan your week around the raw number.
Stop Guessing. Start Scoring.
Lyre Leads searches Google Maps, enriches every result with 50+ data points, and scores each lead with AI. Find the businesses that actually need your help. Free plan includes 500 tokens to try it out.
Start free, no credit card required
Lyre Leads