How to Use AI for Private Equity Market Mapping

A market map is easy to start and hard to finish.

You begin with five familiar companies. A database adds 40. Search finds another 20. By Friday afternoon, the file has 113 rows, three definitions of the market and no good answer when someone asks if the best target is missing.

AI helps because small companies describe themselves in inconsistent ways. It can generate adjacent search terms, recognise similar business descriptions and clean up messy company data. That is useful in fragmented sectors, where the strongest target may never use the label in the investment thesis.

The market boundary determines the quality of the list. A loose definition simply produces a larger pile of weak names.

Write the inclusion rule first

Put the market definition in plain English. It should be specific enough that two associates classify the same company the same way.

“European testing businesses” is too broad. A workable version might read:

Independent providers of recurring, regulation-driven testing and inspection services to food manufacturers in the UK and Ireland, excluding equipment makers, in-house laboratories and generalist consultancies.

Before searching, add four fields to the file:

  • Must-have characteristics
  • Explicit exclusions
  • Geography
  • A size proxy, such as employees, locations or estimated revenue

Save the definition and date it. The thesis will change as the team learns, and the file should show when that happened. Otherwise, an old list quietly becomes a new list halfway through the work.

Set up the harness

A market map usually becomes a folder rather than a single chat. Put the thesis, seeds, raw discoveries and reviewed output in separate files:

/Market Map - UK Fire Safety
  market_definition.md
  seeds.csv
  raw_candidates.csv
  enriched_candidates.csv
  exclusions.csv
  search_log.md

Claude Code or Codex can work well here because both can operate across a local directory and run small scripts. That makes deduplication, domain matching and CSV checks repeatable. Cowork and ChatGPT Work are useful when the inputs are documents, browser research and finished tables. Work can use local folders in the desktop app when that permission is available. Cowork can work with local files through Claude Desktop. If you use a connector for Drive or another data source, remember that the connector inherits your permissions and may run through the provider’s cloud service.

Keep one human-controlled file: market_definition.md. The agent may suggest edits, but it should not quietly move the boundary while classifying companies.

A worked discovery and enrichment prompt

Take a UK fire-safety services thesis. The inclusion rule is recurring inspection, testing, maintenance or compliance services for commercial customers. Equipment-only distributors, installers with no service base and generalist facilities managers are excluded.

Put 10 known good seeds in seeds.csv, with a sentence explaining why each fits. Then give the agent this prompt:

Read market_definition.md and seeds.csv. Build a search plan before collecting companies. Create query groups for service synonyms, regulations and certifications, customer problems, trade bodies, local terminology and adjacent services. Write the groups to search_log.md. For each group, state which part of the inclusion rule it tests. Do not change the market definition.

Use the approved research sources to collect candidate companies. Append them to raw_candidates.csv with company_name, website, headquarters, discovery_source, discovery_url, query_group, date_accessed and one exact quote from the source. Keep duplicates. Do not classify a company from a search-result snippet alone.

The next prompt turns the raw list into a review queue:

Enrich each row in raw_candidates.csv from the company’s own website and one independent source where available. Create enriched_candidates.csv with standardised_name, domain, locations, service_lines, customer_types, regulatory_markers, recurring_service_evidence, estimated_size_proxy, fit (include/exclude/review), fit_reason, source_urls and evidence_quotes. Apply only the rules in market_definition.md. Use REVIEW when the evidence is incomplete. Never infer recurring revenue from the words maintenance, compliance or subscription without a source quote.

For one candidate, a defensible row might read:

Company: Northgate Fire Compliance
Fit: REVIEW
Reason: offers annual fire-alarm inspection and extinguisher servicing to commercial sites; ownership and scale unclear
Evidence: "planned inspection and maintenance contracts" - company services page
Size proxy: 3 service locations; employee count not verified
Sources: company services page; BAFE register entry

That is enough to decide the next research step. It is not enough to put the company on an IC slide as a verified target.

Run a sanity check before calling the map complete

Start with 20 rows: five includes, five excludes, five reviews and five duplicates or edge cases. Re-run the classification manually from the linked evidence. Track false positives and false negatives by query group. A broad query that adds 30 names and produces one qualified company may still be useful, though the file should show the cost.

Then test coverage from the other direction:

Compare enriched_candidates.csv with the membership lists or registers in search_log.md. Report known members or certified providers absent from our file. Group misses by likely cause: terminology, geography, domain matching, source access or classification. Suggest one additional query for each cause. Do not add any company to the final map until a source page has been opened and quoted.

I also check the five seed companies. If the process cannot rediscover most of them without feeding their names into the search, the search language is too narrow. Finally, take 10 excluded companies and ask whether the exclusion reason still follows the written rule. This catches thesis drift faster than reading another 100 rows.

Use a seed list

A small set of good examples teaches the model more than a page of abstract instructions. Seeds can come from known deals, trade associations, conference exhibitors, expert calls, banker materials and the team’s network.

Ten well-chosen companies will usually beat 100 weak examples. Record why each seed fits: end market, service model, customer base, regulatory driver or some combination of those.

This also exposes false friends. Two businesses may share an industry label while having different economics. One sells equipment and the other provides the recurring service around it. A keyword screen will mix them together unless the distinction is explicit.

Expand the search language

Ask for the terms a company could use without using your preferred category. Useful groups include:

  • Synonyms and older industry terms
  • Customer-problem language
  • Regulatory standards and certifications
  • Trade-association categories
  • Product and service descriptions
  • Local-language terms
  • Adjacent services commonly offered by the same company

Run the terms as separate searches. One giant query hides which route found each company and makes it difficult to see where the noise came from.

A fire-safety thesis, for example, may require searches around inspection, compliance, testing, suppression systems, alarm maintenance and local regulation. The best small operator will not call itself a “fire-safety platform”. Founders do not tend to write like IC papers.

Combine discovery routes

No single database has the full market. Search engines do not either. A defensible map draws from several places:

  1. Company and transaction databases
  2. Trade bodies and member directories
  3. Conference exhibitor and sponsor lists
  4. Regulator and certification registers
  5. Local search results and maps
  6. Competitor, supplier and customer pages
  7. Job postings
  8. Deal announcements and old banker materials

AI can extract names and standardise the entries. Keep the discovery source. A company found in a regulator’s register carries different evidence from one listed in a generic directory.

For the human side of origination, see The Art of Deal Sourcing in Middle Market Private Equity.

Build the file around evidence

My usual columns are:

  • Company and website
  • Headquarters and operating geographies
  • Products and services
  • End markets and customer types
  • Ownership and sponsor history
  • Estimated scale and the basis for it
  • Locations and employee count
  • Relevant certifications
  • Reason for fit
  • Main concern
  • Source URL and date checked
  • Confidence level
  • Next research action

The final four columns prevent a good description from hardening into fact after several rounds of copying between spreadsheets and slides.

Let AI populate a first pass. Any field that affects the shortlist should have a source. “Estimated revenue: $25m” without a source or method adds very little.

Score after you understand the market

Early scoring creates false precision. Build enough of the map to see the range of business models, then set the scorecard.

Five or six factors are normally enough:

  • Strategic fit
  • Revenue quality
  • Market position
  • Size
  • Geography
  • Ownership and likelihood of transacting

Facts and judgments should sit in different fields. Employee count is a fact, subject to source quality. “Attractive founder succession situation” is a judgment, often based on thin evidence.

I prefer red, amber and green with a short explanation. A 73.4-point score looks scientific and rarely survives a real discussion. If the rating cannot be explained in one sentence, the deal team will ignore it.

For bolt-ons, connect the scorecard to the platform plan. The guide to bolt-on acquisitions in private equity covers why integration fit and management capacity can matter as much as the entry multiple.

Look for companies the process missed

Teams naturally review false positives. The bigger risk is the good company that never entered the file.

Run coverage checks:

  • Which regions have fewer companies than expected?
  • Which trade-body members are absent?
  • Which names appear on competitor or customer pages but nowhere in the map?
  • Which search terms produced no results and may need local terminology?
  • Which subsectors have no representative company?
  • Which acquired businesses disappeared into larger groups and stopped ranking independently?

The model can review gaps against the agreed definition. A useful instruction is:

Review the current list against the inclusion and exclusion rules. Identify coverage gaps by geography, service line, certification and customer segment. Only propose a company when a live source shows why it may fit.

That final constraint cuts out much of the rubbish.

Verify the shortlist

A website shows how a company describes itself. It does not prove ownership, scale or performance. Check the original source for every field that changes the shortlist.

Ownership needs particular care. Old articles, database records and group websites often conflict. Store the claim with an as-of date. Do the same for headcount, locations and acquisition history.

Before a company moves from long list to shortlist, a reviewer should open the sources, confirm the business model and write one sentence on fit and one on the main concern. This small bit of friction catches most hallucinations without sending the whole exercise back to manual research.

Refresh rather than rebuild

Market maps decay. Companies are acquired, founders retire, websites change and businesses move into adjacent services.

Add a last-checked date and refresh status. When the theme comes up again, update stale sources and rerun the discovery routes that need it. New names can then be assessed against the same schema.

The first map may take weeks. The third refresh should take much less time.

In practice

Use AI to widen the search language, expand a seed list and organise the evidence. Define the boundary before starting, keep sources beside the claims, and run a deliberate search for missed companies. The finished map should stand up when the deal team asks why this is the investable universe.

Tool references


Tags


Related Articles

{"email":"Email address invalid","url":"Website address invalid","required":"Required field missing"}
>