arrow_back

Beyond Vibe Coding Vol. 4 — Professional Skepticism in the Age of AI

AI has knowledge. But Professional Skepticism is still human.

Through Vol. 3, I wrote primarily from the perspective of a specialist on the systems development side — focusing on the design and accountability of building with AI.

This time, I want to reflect on the risks of using AI, as a Certified Public Accountant, drawing from something that happened recently.

A note on scope: most of my readers have a background in accounting, so I'll proceed without explaining every technical term.
I'll also use abbreviated forms for some longer terms. Writing out "cash and cash equivalents" every time gets unwieldy.


I Asked AI to Create a Diagram of IFRS 18's Impact

I was working on a diagram to illustrate how IFRS 18 affects the presentation of the cash flow statement.

https://ifrs-labo.com/posts/ias-7-ifrs-18-biggest-cash-flow-reform-in-30-years

Specifically, three changes:

A simple Before/After comparison — nothing more. I had already finished writing the article explaining the changes, and then asked AI to produce the diagram.


Take 1: It Looked Good

Here's what was generated first.

 Can you spot the critical error?

It looked polished. IAS 7 current practice on the left, IFRS 18 applied on the right, key changes in the center. Arrows showing the correspondence, color-coding that made sense.

But I immediately noticed something that would make it impossible to explain.

Simply applying IFRS 18 had changed the ending cash balance.

This is exactly the kind of discrepancy that makes explanation impossible.

The ending balance in the Before column: 620. In the After column: 420.

IFRS 18 is a standard that changes the presentation and classification of flow information. It has absolutely no effect on the actual movement of cash.

Whichever presentation method you use, the ending cash balance should be the same — 620.

This is a conclusion you can reach without knowing the fine details of IFRS 18.

(For what it's worth, I actually noticed first that the opening operating profit figure didn't match the theoretical value — and was then dismayed to find the ending cash balance was also wrong.)


Differences Between Models

I was curious. Was this a problem specific to one model, or something more general across AI?

So I tested the same diagram across multiple models — GPT, Gemini, Claude, and others.

What I found was this: whether the error was caught depended not on the tool or model, but on how the question was asked.

Why?

AI does not possess the understanding that "in this context, this number must not change" — the kind of professional skepticism that arises from accounting judgment. More precisely: it doesn't know where the threshold is.

A human professional who has a duty to explain will feel it immediately upon seeing the diagram — "Wait, I can't explain this." IFRS 18 is a reform of flow presentation, not a redefinition of cash. Changing that figure would mislead. That foundational sense of responsibility triggers the alert.

AI still struggles to implicitly infer "the level of contradiction that would make explanation impossible."


I Pointed It Out. Something Else Broke.

When I flagged the error in the ending balance and the incorrect opening operating profit figure, a revised version appeared — Take 2.

 The ending cash balance now matches.

The balance was back to 620. At first glance, it appeared fixed.

 You can check this with mental arithmetic.

But when I checked the subtotal for financing activities — the total came to (20).

The correct figure is (220).

This is a different category of problem from understanding the presentation standard. This is arithmetic.

A cash flow statement with a column that doesn't add up should never be presented. That's a matter of professional pride.

AI doesn't know the feeling of breaking into a cold sweat when a client catches a contradiction or a simple mistake.

At this stage, AI cannot reliably guarantee numerical consistency. Producing output takes priority.

To its credit, AI is highly committed to producing an answer.— but "I'm not confident, so I won't output anything" is not in its vocabulary.

That's simply a characteristic of the tool. Understanding that characteristic is what allows you to use it effectively.


What AI Struggles With Is Not Explaining Knowledge

What struck me this time is that AI's weakness is not explaining knowledge.

The major models have learned the content of IFRS 18 reasonably well. Ask them to explain it, and they'll answer accurately.

The problem is that context-dependent skepticism — "something feels off about this number in this context" — does not activate.

In accounting and auditing, this stance is called professional skepticism: not accepting information at face value, but continually asking, "Is this really correct?"

It's the same thing I wrote about in Vol. 2 — the habit of repeatedly asking "why?"

AI answers when asked. But it takes no responsibility for its output.

Whether that output is fit for purpose, whether it's something you can present with confidence — that judgment still belongs to humans.

And precisely because the output looks polished, the risk of missing a critical error and using it as-is is higher than ever.

This is the version I ultimately used.


What Will Be Required in the AI Era

"You shouldn't use AI-generated materials as-is" — I hear this often. But what I felt this time was something more acute than that standard warning.

This time I was writing an article, so I had the time to review carefully. But in a high-speed closing environment — would you actually catch an error at this level without missing it? That's a genuinely concerning thought.

The majority of what AI generates is correct.

It looks well-organized. The explanations sound right. And that is precisely why the few percent of critical errors hidden within are harder to find than ever before.

What you need to catch those errors is not memorized knowledge.

"This number shouldn't be changing." "Something feels off about this explanation."

It's the capacity to sense that something is wrong.

The reason I noticed the problem in the AI-generated diagram wasn't that I had memorized the text of IFRS 18.

It was because I understood the underlying premise: since this is a change in presentation classification, the cash balance itself cannot change.

At IFRS-LABO, I've consistently advocated for understanding the substance of standards rather than memorizing them — because genuine understanding lets you apply the principles to new standards, and lets you evaluate whether AI output is right or wrong.

That conviction feels even more important now, in the AI era.

AI has knowledge.

But whether its conclusions are truly sound — that judgment still belongs to humans.

This kind of capacity cannot be built by memorizing rules. It is forged only through real-world experience and repeated, deliberate practice.

That is precisely why the habit of always asking "why does it work this way?" matters.

Professional skepticism is not developed by reading standards.
It is developed by repeatedly finding inconsistencies and explaining them.

IFRS-OneQ was built for that kind of training.

To understand the substance. To sense the discrepancy. To verify it.

That is the kind of professional who will hold their value in the AI era.