arrow_back

AI vs. Technical Research — Episode 2: Teaching AI Where to Look

How missed evidence became the starting point for a better research process

In Episode 1, I wrote about a problem that is easy to underestimate: AI can give an incomplete answer that sounds complete.

While researching companies that had early adopted IFRS 18, several AI models suggested that very few real-world cases existed. Their answers sounded reasonable. The standard was new, early adoption was optional, and relevant disclosures were often buried inside company reports.

But a manual search revealed a case that AI had missed. Further research uncovered more.

My first reaction could have been:

AI is unreliable for this kind of research. I should return entirely to manual search.

But that conclusion did not feel right either.

AI had helped me understand unfamiliar terminology, generate possible search directions and cover more ground quickly. The problem was not that AI had no value. The problem was that I had given it the wrong role.

I had asked AI to complete the investigation.

So I changed the assignment.

Instead of asking AI for the final answer, I began giving it confirmed evidence, challenging what it had missed and using each discovery to decide where we should search next.

That was when AI became more useful.

From one question to several research tasks

My original question had sounded simple:

Which companies have early adopted IFRS 18?

Looking back, that single question contained several different tasks.

AI needed to:

That was too much responsibility to place on one broad prompt.

I had treated AI like a senior researcher who already understood the scope, terminology and evidence requirements. In reality, I had not yet defined those requirements clearly enough.

The first improvement was therefore not a more sophisticated prompt. It was breaking the investigation into smaller decisions.

First, find possible cases. Then inspect the underlying documents. Then decide whether each company actually applied IFRS 18. Finally, record the evidence and any uncertainties.

This separation became fundamental to the rest of the research.

A missed case became a useful clue

When I manually found a case that AI had missed, I did not simply add the company to the tracker and move on.

I brought the case back to AI.

I asked why it might have been missed. I showed it the language used in the report and asked for alternative search terms. I used the company’s jurisdiction, reporting period and document type to generate new directions.

The missed case began to reveal several possible weaknesses in the original search:

Each weakness suggested a different search.

Instead of repeatedly asking, “Find more IFRS 18 early adopters,” I could search around language such as:

I could also search by jurisdiction, reporting period, exchange filing, interim-report type or a list of companies reporting under IFRS.

The failed search had produced something valuable: information about how to make the next search better.

A case that AI misses is not only a correction. It can become evidence about where the research process is weak.

Building a findings database

As the number of searches increased, I needed a way to prevent the investigation from becoming a collection of disconnected links and conversations.

So I began building a findings database.

For each potential case, I recorded information such as:

The conclusion field was especially important. A company could be marked as confirmed, rejected or still unresolved.

Rejected cases were not deleted.

For example, a report might mention IFRS 18 in its section on new standards but state that the company had not yet adopted it. That was not an early-adoption case, but it remained useful evidence. Recording the reason for rejection prevented the same company from returning later as an apparently new lead.

The database also made patterns visible.

Some jurisdictions produced reviewed interim financial statements relatively quickly. In other markets, the available evidence consisted mainly of future-adoption disclosures. Some companies used IFRS 18 in voluntary reference information, while others applied it within statutory financial statements.

These differences were difficult to see when each search result was considered separately.

My background is in product, GTM, marketing and AI rather than accounting. Perhaps because of that, I began treating the investigation in a way that felt similar to product discovery: collect evidence, identify patterns, test assumptions and use what we learn to improve the next question.

The database became the shared memory of the research process. AI could help explore it, but the evidence did not depend on AI remembering every previous conversation correctly.

Accounting expertise changed where we searched

Organizing the evidence improved the process, but it also exposed a limitation in my own perspective.

I could identify search patterns, compare language and structure the findings. But understanding why evidence appeared in one market and not another required accounting and regulatory context.

This was where discussion with my CPA co-builder became essential.

Our conversations raised questions that were not obvious from a simple keyword search:

These questions did not merely help us evaluate results after they appeared. They changed where we searched in the first place.

For example, a lack of visible cases in one country did not necessarily mean that companies were less prepared. It could reflect reporting deadlines, endorsement status or the absence of reviewed interim reporting.

Similarly, finding an IFRS 18-style presentation did not automatically prove formal adoption. The applicable reporting framework and basis of preparation still mattered.

AI could generate more keyword combinations. Accounting expertise helped define which combinations were meaningful.

That distinction changed my understanding of human involvement in AI-assisted research.

The human expert is not only there to check the answer at the end. Professional knowledge helps define the search space before the answer is produced.

The research became a loop

Over time, the workflow became iterative:

  1. Use AI for orientation.
    Identify possible terminology, jurisdictions, companies and document types.

  2. Search manually.
    Try different keyword forms, local-language terms, reporting periods and primary-source websites.

  3. Verify the evidence.
    Open the underlying company report and read the relevant disclosure in context.

  4. Record the result.
    Add confirmed, rejected and unresolved cases to the database.

  5. Discuss the pattern with domain experts.
    Use accounting expertise to interpret why the evidence appeared in that form.

  6. Return the findings to AI.
    Ask what was missed, what related searches are possible and where similar evidence might exist.

  7. Start the next round.

No single step was sufficient on its own.

AI expanded the possible search directions. Manual search exposed gaps. Primary documents supplied evidence. Professional knowledge added meaning. The database preserved what we had learned.

The improvement did not come from discovering one “perfect prompt.” It came from repeated interaction between evidence, judgment and better questions.

Sometimes AI found what I could not

An important part of this experience is that manual search did not always perform better.

After several rounds of feedback and additional context, AI occasionally found company reports that were difficult for me to reproduce through a conventional search.

The relevant document might have been poorly indexed. The disclosure might have appeared deep inside a PDF. The company might have used unexpected wording, or the result might have required connecting several search concepts.

This prevented me from reaching another false conclusion: that the safest solution was simply to replace AI with Google and manual research.

Neither method was consistently superior.

They had different strengths.

AI was useful for:

Human researchers remained essential for:

The goal was no longer to decide whether AI or manual research was better.

The goal was to design a process in which each could compensate for the limitations of the other.

How I now use AI for technical research

This experience led me to adopt several practical rules.

1. Define the evidence threshold before searching

What will count as confirmation?

For our IFRS 18 tracker, an AI answer, search snippet or secondary article could identify a potential case. It could not confirm one. Confirmation required evidence from an official company report or another authoritative primary source.

2. Separate discovery from confirmation

Discovery should be broad. AI and manual search can both generate possible cases.

Confirmation should be narrow. Only evidence that meets the predefined standard should enter the final tracker.

Keeping these stages separate reduces the risk that a convincing lead silently becomes an accepted fact.

3. Feed failures back into the process

When AI misses something, ask what the missed case reveals about terminology, jurisdiction, source type or timing.

The objective is not merely to correct the answer. It is to improve the next search.

4. Record negative findings

A rejected case can be just as useful as a confirmed one.

It documents what was checked, why it did not qualify and which misleading signals might cause it to reappear.

5. Use professional expertise before and after the search

Expert judgment should not be reserved for final review.

It can help decide which markets to examine, which documents matter, which terminology is meaningful and which apparent differences require further investigation.

The lesson from the second experiment

Episode 1 taught me that agreement among AI models does not prove that an answer is complete.

Episode 2 taught me something equally important:

AI research improves through iteration, not through one perfect prompt. Its value increases when evidence and professional knowledge define the search space.

AI became a much better research partner when I stopped asking it to own the entire conclusion.

But a better process did not make every result reliable.

After repeated feedback and refinement, ChatGPT offered to monitor new IFRS 18 early-adoption cases automatically. The idea sounded useful: instead of repeating the search manually, the system would notify me when it found a new adopter.

Then one day, it sent exactly the alert I had asked for.

It identified a company, gave a specific adoption date and described the accounting treatment. The finding looked detailed and well supported.

There was only one problem:

The company had not adopted IFRS 18.

Next: Episode 3 — The “Confirmed” Case That Wasn’t