AI for CIM Review: A Private Equity Due Diligence Workflow

Most people start by uploading the CIM and asking for a summary. That may save half an hour, although it rarely improves the analysis. The CIM is already a summary, written to sell the company.

The useful questions sit underneath it. Which claims matter? Which pages disagree? What has to be true for the forecast to work? What evidence is missing?

AI can help answer those questions. I use it for extraction, reconciliation and challenge. The deal team still makes the judgment.

Clear the security gate first

A CIM can contain customer names, employee data, pricing, forecasts and material covered by an NDA. Check that the firm’s approved environment can accept the document before uploading it. A consumer chatbot should never become a shadow data room.

If there is no approved environment, build and test the workflow on public filings or a blank example. Removing the company name from a live CIM does not make the rest of the document anonymous.

This rule is easy to follow when the timetable is relaxed. It matters most when the team is rushing.

Choose the working environment

The tool matters less than the access pattern. I use a document-oriented agent when the job starts with a CIM and ends with a workbook or a set of diligence questions. Claude Cowork can work with a folder on the computer through the Claude desktop app. ChatGPT Work can also open a local folder in the desktop app when local access is enabled. Both suit a contained review where the source files and outputs live together.

Claude Code and Codex make sense when the deal folder already has a repeatable structure. They can work through a local directory and run scripts, which is useful for checking tables, normalising units and writing outputs to CSV or Excel. I normally create a clean subfolder first:

/Deal Alpha
  /source/CIM.pdf
  /source/financial_appendix.xlsx
  /work/cim_extract.csv
  /work/cim_questions.md
  /work/verification_log.md

Give the agent access to /Deal Alpha, rather than the wider PE directory. Keep write access inside /work. Add a short standing instruction in CLAUDE.md for Claude Code or AGENTS.md for Codex: preserve page references, never overwrite source files, write NOT FOUND when evidence is absent, and keep calculations separate from extracted facts.

Cloud features and connectors need the same approval as an upload. A locally selected folder does not settle the data-policy question by itself. Check where the session runs, what is retained and which connectors are enabled.

A worked CIM extraction

Assume the CIM covers a testing business with FY2023 to FY2025 financials, a 2026 budget and a 2027 to 2030 forecast. Start with a narrow job. This is the first prompt I use in Cowork, Work, Claude Code or Codex:

Review source/CIM.pdf and source/financial_appendix.xlsx. Create work/cim_extract.csv with one row per fact. Use these columns: topic, metric, value, unit, period, status (historical/budget/forecast), definition, source_file, page_or_tab, exact_source_text, confidence, reviewer_note. Extract revenue, gross profit, adjusted EBITDA, reported EBITDA, capex, net working capital, customer concentration, retention, price, volume and headcount. Preserve the units shown in the source. Do not calculate missing figures. If a requested item is absent, add a row with value NOT FOUND. If the PDF and workbook disagree, keep both rows and mark the conflict in reviewer_note. Finish with a short list of pages or tabs that could not be read cleanly.

Then run a second pass with a different job:

Read the completed work/cim_extract.csv against the source files. Find: (1) same metric and period with different values; (2) reported and adjusted figures sharing an unclear label; (3) totals that do not reconcile with components; (4) chart values that appear to use different units; and (5) forecast claims with no visible operating bridge. Write every issue to work/cim_questions.md. For each issue, show both citations and the smallest management question that would close it. Do not rewrite or correct the extracted values.

That produces a useful first cut. It still needs a reviewer.

Verify the output in a fixed sample

I use a mechanical check before reading the conclusions. Pick 15 rows across the file: five financial figures, five operating metrics and five awkward items such as footnotes, chart labels or add-backs. Open the cited page or tab and check the value, period, unit, definition and exact quote. Record each result in verification_log.md.

A second prompt can find likely errors, although it cannot certify its own work:

Audit work/cim_extract.csv. Re-open every cited page or tab for rows where confidence is below high, reviewer_note is populated, the metric contains EBITDA, retention or concentration, or the period is LTM. Add four columns: citation_opens, value_matches, unit_matches, definition_matches. Use TRUE, FALSE or UNCLEAR. For every FALSE or UNCLEAR result, copy the source wording and explain the mismatch in no more than two sentences. Do not change the original extraction.

If three or more of the 15 sampled rows fail, widen the review. If a key figure fails, check the whole metric family. EBITDA, net debt and customer retention are poor places to extrapolate from a small sample.

Build the extraction table

Open with a fixed table rather than a broad question about the deal. The fields will change by sector, though a starting version should cover:

  • Historical revenue, gross profit, EBITDA and cash conversion
  • Revenue by product, customer, geography and channel
  • Customer concentration, retention and cohorts
  • Price, volume and mix bridges
  • Recurring, project and transactional revenue
  • Normalisations and add-backs
  • Working capital seasonality and capex
  • Net debt, leases, pensions and other debt-like items
  • The forecast and its main assumptions
  • Market statistics and their original sources

For every entry, capture the page, period, unit and wording used in the CIM. Missing data should be marked “not found”. Estimates and inferences belong in separate columns.

This split works well. The machine handles the clerical sweep; the associate chooses the fields and checks that the output makes sense.

Review the CIM in four passes

Pass 1: Extract the facts

Populate the table. Keep history apart from budget and forecast. Keep reported figures apart from adjusted figures. Preserve units.

Small presentation details cause plenty of mistakes. A chart labelled in thousands gets copied into a table labelled in millions. A last-twelve-month figure lands beside calendar-year numbers. A percentage is mistaken for a percentage-point change. This pass should catch those errors before anyone writes a view.

Pass 2: Find internal inconsistencies

Ask for contradictions and unexplained changes, with page references. Common examples include:

  • The overview calls revenue recurring while the customer pages show frequent churn
  • A “one-off” EBITDA adjustment appears in several years
  • The market section claims low cyclicality while monthly trading swings sharply
  • The customer slide and financial appendix show different concentration figures
  • The forecast assumes margin expansion without a headcount or pricing bridge

Some discrepancies will have simple explanations. Put them on the question list anyway. It is better to close a point deliberately than lose it somewhere in 120 pages.

Pass 3: Reverse-engineer the forecast

Translate the plan into operating assumptions. How many customers must be added? At what price? What sales productivity is implied? How much gross-margin improvement is required? How much cash will working capital absorb?

Compare those assumptions with the historical record. If the company has never grown faster than 8%, an 18% forecast needs evidence beyond a blue line sloping upwards.

Pass 4: Build the missing-evidence list

Every unsupported claim should become an information request, a management question or a diligence workstream. This turns the first review into something the wider team can use.

Spend more time on the weak spots

Quality of earnings

Create a schedule of adjustments showing amount, period, rationale, recurrence and supporting evidence. “Non-recurring” is an argument rather than an accounting category. Repeated restructuring, recruitment or professional fees should be challenged.

Revenue quality

Split price from volume, organic from acquired growth, and new customers from expansion. Check whether retention means logo retention, gross revenue retention or net revenue retention. Those measures answer different questions.

Customer concentration

Look beyond the top-ten table. The same customer may appear through several legal entities. Distributors can hide end-customer exposure. Concentration may also rise in the forecast. A 15% customer whose contract expires three months after closing deserves more attention than the headline percentage suggests.

Cash conversion and debt

EBITDA gets the attention; cash pays the debt. Pull out working capital, maintenance capex, leases, capitalised development costs, tax leakage and other debt-like obligations. This is often where an attractive entry multiple begins to look less attractive.

For the wider set of workstreams, see Types of Private Equity Due Diligence.

Make the prompts narrow

Broad prompts produce broad answers. A better instruction is:

Create a table of every EBITDA adjustment in the document. Include the amount, period, stated rationale and page number. Flag adjustments that appear in more than one period or are described inconsistently. Do not assess validity. If a field is absent, write “not found”.

Follow that with:

Using only the extracted table and cited pages, list the five strongest challenges to adjusted EBITDA. For each one, give the evidence, the management question and the document needed to close the point.

The first step constrains the evidence. The second turns it into a review list. Without that sequence, the output tends to be a set of risks that could apply to almost any company.

Reconcile later material

The CIM is the opening statement. As the model, quality-of-earnings report and data room arrive, update the extraction table and log differences:

  • CIM claim
  • Later evidence
  • Difference
  • Explanation
  • Status
  • Owner

This log is worth keeping. A week-one question may receive a verbal answer in week two, be contradicted by a spreadsheet in week three and remain marked “resolved” in the final memo.

AI is useful for comparing versions and spotting changes. Materiality calls need a person.

Keep the source beside the claim

Page references are the minimum for CIM work. Later entries can include a workbook tab and cell range, data-room file name, meeting date or diligence-report section.

Another team member should be able to reproduce the point without asking where it came from. If they cannot, the output is prose rather than analysis.

Where the tool stops

AI cannot decide if the deal is good. It cannot approve an add-back, select the downside case or judge management credibility. It also cannot know when a perfectly extracted figure rests on a poor definition. “Recurring revenue” might mean contracted revenue in one company and revenue that happened twice in another.

The best use is a faster first pass and a sharper challenge list. The investor still has to decide what is believable.

In practice

Break the CIM review into controlled jobs: extract, cite, reconcile, challenge and turn the gaps into questions. The result should be an evidence-backed map of what the team knows, what remains uncertain and what must be proved before IC.

Tool references


Tags


Related Articles

{"email":"Email address invalid","url":"Website address invalid","required":"Required field missing"}
>