arrow_back

A Small Feature, and a Familiar Way of Working

Building the new Mock Exam for IFRS-OneQ with AI

We recently added a new Mock Exam to IFRS-OneQ. It is a simple assessment feature: 10 questions in 15 minutes, covering the six proficiency areas used in IFRS-OneQ.

Unlike the normal learning modes, the Mock Exam is designed specifically for assessment. Explanations are not shown during the exam, and the results are kept separate from normal learning progress and Skill Proficiency. Each attempt is recorded separately, so users can look back at their previous scores and completion times.

The purpose is straightforward. IFRS-OneQ already provides a place to learn and practise IFRS. The Mock Exam gives users a place to step away from that process and ask a different question: How much do I actually know right now?

It is a relatively small feature, but it is also a deliberate step toward where I would like to take IFRS-OneQ next.

Why a Mock Exam?

We started by looking more broadly at what users may ultimately want from IFRS learning. Our current hypothesis is that much of the value lies beyond learning itself. Users may want to become better at their work, demonstrate what they know, or eventually use that knowledge to create new career opportunities.

As part of that process, we looked at competing learning products, professional exams and certification models. One direction we would eventually like to explore for IFRS-OneQ is a Certificate, so that users could potentially demonstrate what they have learned rather than simply complete questions.

That raises a difficult question: how much credibility can an online assessment have without making the user experience unnecessarily heavy? Identity verification, monitoring and other controls may strengthen credibility, but they also change the nature of the product. We do not think that balance has been solved yet.

So we did not jump directly to a Certificate. Instead, we looked for a smaller but meaningful step in the same direction. Separating learning from assessment was useful in its own right and gave us something concrete to build on. That became the Mock Exam.

While product ideas are often discussed within our small team, I handle the actual implementation of IFRS-OneQ myself, working directly with AI.

For the Mock Exam, the whole process — including competitor research, thinking through the product direction, investigating the existing system, defining and reviewing the requirements, and implementing the feature — took roughly half a day.

The speed is not really the interesting part

That development speed did not particularly surprise me. Much of IFRS-OneQ has been built in the same way. Once I decide that something is worth building, the cycle from investigation to implementation can now move extremely quickly.

I currently use Claude mainly for research, planning and review, and Cursor mainly for investigating the existing code and implementation. The process moves back and forth between them. Research affects the product idea, the product idea leads to technical investigation, the investigation changes the plan, and the plan sometimes exposes another decision that needs to be made before implementation continues.

What I find more interesting is what I am doing while all of this is moving so quickly. I am not reviewing every line of code, nor am I independently reproducing everything the AIs investigate. I am mostly looking for something more fundamental: Are they working from the right premises?

During the Mock Exam development, the answer was “not quite” several times.

The important corrections were about premises

During the Mock Exam development, I had to correct the AIs several times. Some were straightforward misunderstandings. At one point, the current number of questions in IFRS-OneQ started to be treated as if it were relatively fixed, even though we continuously add new questions. At another point, the discussion proceeded on the assumption that the Mock Exam would be freely available, which was not the product model I had in mind.

These were useful corrections, but they were essentially misunderstandings about the product. The timer led to a more interesting case because the AI's proposal itself was perfectly reasonable.

The AI suggested adding a Pause function. For a timed online exam, that makes intuitive sense: if a user is interrupted, stop the timer and let them continue later. Technically and from a general UX perspective, there was nothing obviously wrong with the idea.

But instead of asking how we should implement Pause, I went back to what I wanted the Mock Exam to do.

One purpose of the Mock Exam is to help users check their answering pace under time pressure. In a real exam, the clock does not stop when you are interrupted or need more time. I therefore did not want the Mock Exam timer to stop either. The question became not “How should Pause work?”, but “How can we make a continuous timer useful in the real lives of the people using IFRS-OneQ?”

That thinking also influenced the format itself. IFRS-OneQ is primarily designed for busy professionals, so I wanted each attempt to fit into a relatively small part of the day. At the same time, 10 questions are enough to give users an early indication of their pace. The idea behind 10 questions in 15 minutes is that if users can work through this short set comfortably within the time limit, they should be building enough pace to leave themselves some room in a longer real exam.

A short format also means that we do not need to solve every user need by making a single attempt more flexible. If someone wants more assessment, they can simply take another attempt. If they need to stop early, they can use “Finish here” to end the attempt at that point, record the result and move on.

So I chose Finish here rather than Pause. Not because Pause is inherently a bad feature, and not because the AI's suggestion was technically wrong. It was simply not the choice I wanted for this particular product.

This distinction is important to how I work with AI. Sometimes I am correcting a factual misunderstanding. But at other times, AI has understood the problem perfectly well and proposed a completely reasonable solution — and I still choose something else because I am working from a different view of the user, the purpose of the feature, or the behaviour I want the product to encourage.

If I don't understand why, I don't move on

This is probably the part of AI-assisted development where I spend most of my own attention. If an AI proposes something and I do not understand why it is necessary, I ask. If the answer depends on the existing system, I have the relevant part investigated. If what initially looks like a technical requirement turns out to be a product choice, then I make that choice.

The important point is not that my initial view must always win. Sometimes further investigation shows that the AI's proposal makes more sense, and I change my mind. What matters is that an important decision does not quietly become a requirement simply because an AI proposed it.

I do not need to understand every technical detail at the same depth as an engineer, nor do I want to reproduce all of the work that AI has already done. But I do need to understand the important assumptions and decisions well enough to know what I am accepting.

This becomes particularly important when multiple AIs are involved. One AI can research, another can plan, another can review, and another can implement. That can make the process extraordinarily fast, but multiple AIs do not necessarily protect against a shared premise or an unexamined product decision. Once something enters the plan as a “requirement,” every AI downstream can execute it perfectly.

Someone still has to ask whether it should have been a requirement in the first place.

This feels a lot like audit

At some point during this development, I realised that this part of the work felt very familiar to me. It reminded me of audit.

An auditor does not independently recreate everything a company has done. You need to understand enough of the business, processes, systems, explanations and evidence to recognise when something important does not make sense. A calculation can be correct while the assumption behind it is wrong. Several individual explanations can each sound reasonable while the overall picture does not fit together.

That is when you ask why. Sometimes the explanation resolves the concern. Sometimes it shows that your own understanding was incomplete. And sometimes it exposes a problem that would have been difficult to see by reviewing the individual pieces in isolation.

Looking back, I have done variations of the same thing in project management and product management as well. The subject changed, and the people doing the detailed work changed, but a large part of my role remained similar: understand the objective and the overall picture, identify which assumptions matter, notice when something does not fit, and stop to ask why before allowing the work to continue.

Today, much more of the research, planning and execution can be performed by AI. That makes the cycle dramatically faster. What it does not do is remove the need to understand the premises on which that work is being done.

Back to the Mock Exam

After all of that, what users see is intentionally simple: 10 questions, 15 minutes, six proficiency areas. They can step away from their normal learning process, test their current IFRS knowledge, and come back later to see how their results develop over repeated attempts.

The Mock Exam is useful on its own, but it also gives us a practical first step toward thinking more seriously about assessment and, eventually, how IFRS-OneQ users might demonstrate what they know.

From the product hypothesis and competitor research through to the working feature, the whole process took roughly half a day. That speed is increasingly normal in how I build IFRS-OneQ.

What I find more interesting is that, despite the change in tools, my own way of working has not changed nearly as much. I still spend much of my time looking across the work, checking whether the important premises make sense, and asking “why” when they do not.

The tools doing the work have changed dramatically. The need to understand what the work is actually trying to achieve has not.