In Episode 1, I wrote about a problem that is easy to underestimate: AI can give an incomplete answer that sounds complete.
While researching companies that had early adopted IFRS 18, several AI models suggested that very few real-world cases existed. Their answers sounded reasonable. The standard was new, early adoption was optional, and relevant disclosures were often buried inside company reports.
But a manual search revealed a case that AI had missed. Further research uncovered more.
My first reaction could have been:
AI is unreliable for this kind of research. I should return entirely to manual search.
But that conclusion did not feel right either.
AI had helped me understand unfamiliar terminology, generate possible search directions and cover more ground quickly. The problem was not that AI had no value. The problem was that I had given it the wrong role.
I had asked AI to complete the investigation.
So I changed the assignment.
Instead of asking AI for the final answer, I began giving it confirmed evidence, challenging what it had missed and using each discovery to decide where we should search next.
That was when AI became more useful.

My original question had sounded simple:
Which companies have early adopted IFRS 18?
Looking back, that single question contained several different tasks.
AI needed to:
understand what qualified as early adoption;
find the latest company reports;
recognize different ways of describing adoption;
understand the effect of local reporting frameworks;
distinguish adoption from preparation or planned implementation;
and judge whether the evidence was reliable.
That was too much responsibility to place on one broad prompt.
I had treated AI like a senior researcher who already understood the scope, terminology and evidence requirements. In reality, I had not yet defined those requirements clearly enough.
The first improvement was therefore not a more sophisticated prompt. It was breaking the investigation into smaller decisions.
First, find possible cases. Then inspect the underlying documents. Then decide whether each company actually applied IFRS 18. Finally, record the evidence and any uncertainties.
This separation became fundamental to the rest of the research.
When I manually found a case that AI had missed, I did not simply add the company to the tracker and move on.
I brought the case back to AI.
I asked why it might have been missed. I showed it the language used in the report and asked for alternative search terms. I used the company’s jurisdiction, reporting period and document type to generate new directions.
The missed case began to reveal several possible weaknesses in the original search:
The company might not use the exact phrase “early adoption of IFRS 18.”
The disclosure might appear only in an accounting-policy note.
The document might be an interim report rather than an annual report.
The company might describe the effective date without prominently announcing its decision.
The terminology might vary because of local reporting practices or translation.
Each weakness suggested a different search.
Instead of repeatedly asking, “Find more IFRS 18 early adopters,” I could search around language such as:
“early applied IFRS 18”;
“elected to apply IFRS 18”;
“applied from 1 January”;
“basis of preparation” together with “IFRS 18”;
or equivalent terminology used in local-language filings.
I could also search by jurisdiction, reporting period, exchange filing, interim-report type or a list of companies reporting under IFRS.
The failed search had produced something valuable: information about how to make the next search better.
A case that AI misses is not only a correction. It can become evidence about where the research process is weak.
As the number of searches increased, I needed a way to prevent the investigation from becoming a collection of disconnected links and conversations.
So I began building a findings database.
For each potential case, I recorded information such as:
company;
jurisdiction;
reporting period;
report type;
wording used in the disclosure;
link to the primary document;
audit or review status;
and research conclusion.
The conclusion field was especially important. A company could be marked as confirmed, rejected or still unresolved.
Rejected cases were not deleted.
For example, a report might mention IFRS 18 in its section on new standards but state that the company had not yet adopted it. That was not an early-adoption case, but it remained useful evidence. Recording the reason for rejection prevented the same company from returning later as an apparently new lead.
The database also made patterns visible.
Some jurisdictions produced reviewed interim financial statements relatively quickly. In other markets, the available evidence consisted mainly of future-adoption disclosures. Some companies used IFRS 18 in voluntary reference information, while others applied it within statutory financial statements.
These differences were difficult to see when each search result was considered separately.
My background is in product, GTM, marketing and AI rather than accounting. Perhaps because of that, I began treating the investigation in a way that felt similar to product discovery: collect evidence, identify patterns, test assumptions and use what we learn to improve the next question.
The database became the shared memory of the research process. AI could help explore it, but the evidence did not depend on AI remembering every previous conversation correctly.
Organizing the evidence improved the process, but it also exposed a limitation in my own perspective.
I could identify search patterns, compare language and structure the findings. But understanding why evidence appeared in one market and not another required accounting and regulatory context.
This was where discussion with my CPA co-builder became essential.
Our conversations raised questions that were not obvious from a simple keyword search:
Had IFRS 18 been endorsed in the jurisdiction?
Was IFRS mandatory, optional or incorporated into a local framework?
How frequently were listed companies required to publish interim reports?
Were those interim statements externally reviewed?
Could a company voluntarily present IFRS 18 information before formally applying it in statutory financial statements?
Where would a company normally disclose a change in accounting presentation?
These questions did not merely help us evaluate results after they appeared. They changed where we searched in the first place.
For example, a lack of visible cases in one country did not necessarily mean that companies were less prepared. It could reflect reporting deadlines, endorsement status or the absence of reviewed interim reporting.
Similarly, finding an IFRS 18-style presentation did not automatically prove formal adoption. The applicable reporting framework and basis of preparation still mattered.
AI could generate more keyword combinations. Accounting expertise helped define which combinations were meaningful.
That distinction changed my understanding of human involvement in AI-assisted research.
The human expert is not only there to check the answer at the end. Professional knowledge helps define the search space before the answer is produced.
Over time, the workflow became iterative:
Use AI for orientation.
Identify possible terminology, jurisdictions, companies and document types.
Search manually.
Try different keyword forms, local-language terms, reporting periods and primary-source websites.
Verify the evidence.
Open the underlying company report and read the relevant disclosure in context.
Record the result.
Add confirmed, rejected and unresolved cases to the database.
Discuss the pattern with domain experts.
Use accounting expertise to interpret why the evidence appeared in that form.
Return the findings to AI.
Ask what was missed, what related searches are possible and where similar evidence might exist.
Start the next round.

No single step was sufficient on its own.
AI expanded the possible search directions. Manual search exposed gaps. Primary documents supplied evidence. Professional knowledge added meaning. The database preserved what we had learned.
The improvement did not come from discovering one “perfect prompt.” It came from repeated interaction between evidence, judgment and better questions.
An important part of this experience is that manual search did not always perform better.
After several rounds of feedback and additional context, AI occasionally found company reports that were difficult for me to reproduce through a conventional search.
The relevant document might have been poorly indexed. The disclosure might have appeared deep inside a PDF. The company might have used unexpected wording, or the result might have required connecting several search concepts.
This prevented me from reaching another false conclusion: that the safest solution was simply to replace AI with Google and manual research.
Neither method was consistently superior.
They had different strengths.
AI was useful for:
generating alternative terminology;
exploring adjacent companies and jurisdictions;
identifying possible patterns across reports;
processing long documents;
and expanding the number of directions we could investigate.
Human researchers remained essential for:
sensing when the result was too narrow;
understanding the accounting and regulatory context;
deciding what qualified as early adoption;
evaluating the reliability of sources;
and taking responsibility for the final conclusion.

The goal was no longer to decide whether AI or manual research was better.
The goal was to design a process in which each could compensate for the limitations of the other.
This experience led me to adopt several practical rules.
What will count as confirmation?
For our IFRS 18 tracker, an AI answer, search snippet or secondary article could identify a potential case. It could not confirm one. Confirmation required evidence from an official company report or another authoritative primary source.
Discovery should be broad. AI and manual search can both generate possible cases.
Confirmation should be narrow. Only evidence that meets the predefined standard should enter the final tracker.
Keeping these stages separate reduces the risk that a convincing lead silently becomes an accepted fact.
When AI misses something, ask what the missed case reveals about terminology, jurisdiction, source type or timing.
The objective is not merely to correct the answer. It is to improve the next search.
A rejected case can be just as useful as a confirmed one.
It documents what was checked, why it did not qualify and which misleading signals might cause it to reappear.
Expert judgment should not be reserved for final review.
It can help decide which markets to examine, which documents matter, which terminology is meaningful and which apparent differences require further investigation.
Episode 1 taught me that agreement among AI models does not prove that an answer is complete.
Episode 2 taught me something equally important:
AI research improves through iteration, not through one perfect prompt. Its value increases when evidence and professional knowledge define the search space.
AI became a much better research partner when I stopped asking it to own the entire conclusion.
But a better process did not make every result reliable.
After repeated feedback and refinement, ChatGPT offered to monitor new IFRS 18 early-adoption cases automatically. The idea sounded useful: instead of repeating the search manually, the system would notify me when it found a new adopter.
Then one day, it sent exactly the alert I had asked for.
It identified a company, gave a specific adoption date and described the accounting treatment. The finding looked detailed and well supported.
There was only one problem:
The company had not adopted IFRS 18.
Next: Episode 3 — The “Confirmed” Case That Wasn’t