I’m Ngoc Tran, Co-builder of IFRS Labo. My professional background is not in accounting. My expertise is in go-to-market strategy, marketing, and product development. At IFRS Labo I work alongside accounting professionals to turn technical knowledge into useful content, products, and research.
That means I may approach technical research differently from an accounting professional. An accountant may begin with the requirements of the standard, the reporting framework, or the accounting conclusions. I often begin with a different set of questions: How can we find the relevant information? How do we know whether the search is complete? How can AI help us work faster—and where might it create new risks?
But I believe the underlying problem is the same for anyone using AI in a technical field. Whether the subject is accounting, law, regulation, medicine, or another specialized area, the researcher still has to decide whether the information is current, complete, and supported by reliable evidence.
This series is therefore not a textbook explanation of how AI should be used. It is a record of what happened when I used it in a real IFRS 18 early adoption investigation: what it missed, how the research process improved, and why a convincing answer still had to be challenged.
I started with a question that sounded straightforward:
Which companies have already adopted IFRS 18 early?
IFRS 18 is a new accounting standard, so I did not expect a long list. Still, I wanted to identify real company reports—not general implementation commentary, illustrative financial statements, or announcements about future adoption.
An experienced accounting colleague had already searched and found one case: Sony Financial Group in Japan.
I then asked ChatGPT and Gemini to expand the search. Their answers were broadly consistent: practical early-adoption cases were extremely limited, and there did not appear to be much beyond the examples already identified.

That conclusion sounded reasonable. IFRS 18 is new. Early adoption is optional. Relevant disclosures may be buried inside interim reports rather than announced publicly.
It would have been easy to stop there.
But a relatively simple manual search produced another company: du, the UAE telecommunications company.
That discovery changed the research question.
I was no longer asking only, “Which companies have adopted IFRS 18?” I was also asking:
If AI missed a case that I could find manually, what else was missing from its answer?
The problem was not that AI gave me an obviously absurd response. The problem was that it gave me an incomplete response in the language of a conclusion.
There is an important difference between these two statements:
“I found only one case.”
“There is only one case.”
The first describes the result of a search. The second makes a claim about the whole population.
AI can move from the first to the second too easily—especially when information is recent, niche, inconsistently described, or difficult to index.
The fact that more than one model produced a similar answer made it feel more reliable. In reality, the models may have been working with similar visibility gaps: older indexed material, repeated secondary sources, and search results shaped by the same common terminology.
Agreement among models was not independent verification.
The freshness problem became even clearer in the Sony Financial Group case.
At one point, ChatGPT indicated that Sony Financial Group had not adopted IFRS 18. That answer reflected older information, when the company had not yet done so.

But the company’s more recent reporting showed that it had already applied IFRS 18 in its IFRS-based financial information.
(Reference: Sony Financial Group FY2025 Annual report and our analysis)

The earlier statement may once have been accurate. It was no longer accurate for the reporting period I was researching.
This is a particularly difficult failure mode in technical research. An outdated answer can still look authoritative because:
the company name is correct;
the accounting topic is correct;
the cited historical position may once have been correct; and
the wording contains no obvious sign that the information is stale.
For new standards, regulations, and reporting developments, the date is not background information. It is part of the evidence.
After finding the du case, I returned to manual search with a more skeptical mindset.
I varied the terminology rather than relying only on the exact phrase “early adoption of IFRS 18.” I searched by jurisdiction, company-report type, reporting period, and alternative ways a company might describe application of a new standard. I also went beyond the first page of results and opened the underlying financial reports.
That process revealed additional Japanese cases. (can be found at IFRS 18 early adoption tracker)
The important point was not that manual search always beat AI. It did not. In later stages of the project, AI also surfaced reports that were difficult to reproduce through a conventional search.
The lesson from this first stage was narrower:
When the subject is new, niche, or described inconsistently, AI’s inability to find more evidence is not evidence that more evidence does not exist.
Looking back, the search challenge was not caused by one single weakness.

Current-period interim and annual reports can appear faster than models, indexes, and secondary summaries update.
One report may say “early adopted.” Another may say “applied from 1 January 2025.” A company may describe the change in its basis of preparation without highlighting it in the report title or press release.
The relevant sentence might sit inside an accounting-policy note in a long PDF. It may never appear on a highly ranked webpage.
Local endorsement, reporting practices, interim-reporting requirements, and the language used by companies all affect where evidence appears and how searchable it is.
This meant that a broad question to an AI model—however clearly written—could not substitute for a search strategy.
I now treat a negative AI search result as the beginning of a verification step, not the end of the research.
When AI says that it cannot find additional examples, I ask:
What period does the answer actually cover?
Is the model relying on information published before the latest reporting cycle?
Which wording did it search for?
Could the same fact be disclosed using different accounting or local terminology?
Which jurisdictions and report types were included?
Did the search cover interim reports, annual reports, exchange filings, and IFRS-based reference information?
Is the conclusion based on primary sources?
A summary, search snippet, or advisory article may be useful for discovery, but it does not confirm what a company actually reported.
Can one known missed case improve the next search?
A case that AI failed to find is valuable evidence. It can reveal a missing keyword, source type, jurisdiction, or disclosure pattern.
This approach does not require distrusting every AI output. It requires being precise about what the output represents.
An AI-generated list is a search result. It is not yet a complete population.
My first reaction could have been: AI is not reliable for this kind of research, so I should return entirely to traditional methods.
That would also have been the wrong conclusion.
AI was useful for generating initial directions, comparing terminology, suggesting variations, and helping organize what I found. Its failure showed that I had assigned it the wrong role.
I had asked it to complete the investigation.
What I needed was to make it part of an iterative investigation—one in which human judgment, manual search, domain knowledge, and primary-source review continuously improved the next AI-assisted search.
That became the second experiment.
Next: Teaching AI Where to Look
When researching a new technical development, professional skepticism should increase—not decrease—when AI says there is nothing else to find.
The most dangerous answer is not always one that looks wrong. Sometimes it is an incomplete answer that feels complete enough to end the search.