The Day Our Chatbot Got Smarter and Scored Worse

Blogs

Dr. Khanh Phan

•

TL;DR

A chatbot's "accuracy score" can lie to you, depending on how you write the test and who does the grading. We once swapped in a better model and watched our score go down. The problem was not the chatbot. It was that the AI doing the grading was not smart enough.

Why "it looks fine" is a risky way to test

Most companies that deploy a chatbot hit the same wall. It is surprisingly hard to tell whether a chatbot’s answers are consistently right. 

The numbers show it. In a March 2026 survey by Teikoku Databank, 34.5% of companies said they use generative AI at work, 86.7% of those said it delivers results, and the top concern was accuracy of information at 50.4%. A separate survey of IT and DX staff found that 35.2% named unreliable output, or hallucination, as a key issue, second only to security.

So the pattern is clear. People feel the benefit, but they do not trust its correctness.

We went through the same thing with Kotae AI, a chatbot that answers from a company's own documents: manuals, websites, PDFs, spreadsheets. The closest comparison is a new hire who can look things up fast. Fast, but still a new hire.

Our first testing method was simple. Someone typed a few questions, read the answers, nodded, and said it looked good. Looking back, that is a little embarrassing.

The message that woke us up

One day, a customer wrote to us. "Your chatbot told me the store closed in 2019. It did not close."

The question had a false assumption inside it. The right answer was "I don't see anything about a closure in the documents." Instead, the chatbot invented a reason.

The scary part was not that it was wrong. It was that the wrong answer sounded right. People ignore answers that sound wrong. They believe answers that sound reasonable.

That was when we accepted that "looks good" is not a testing strategy. The chatbot needed a real exam.

What kinds of questions should you test with?

We built what we call a golden dataset: a list of questions where a human has already confirmed the correct answer. The first version was about 40 questions, written by hand in one afternoon. Today, it is several hundred questions across seven types.

Two of those types came straight from that customer message.

  • False premise questions: "Why did the store close in 2019?"

  • Out-of-scope questions: "What is the weather tomorrow?"

For both, the correct answer is a polite refusal. A chatbot that refuses correctly is boring. A chatbot that answers confidently is dangerous. We want boring.

The most important skill a chatbot has is saying "I don't know" when it does not know.

This is also a test you can run as a buyer. Ask it something your documents definitely do not cover. If you get a confident answer, that is a warning sign.

Can you let AI grade AI?

Once the dataset passed 100 questions, nobody wanted to read every answer by hand. So we had a second AI grade the answers. It reads the question, the confirmed answer, and the chatbot's answer, then scores it.

It worked well. So well that we stopped reading the answers ourselves.

Then came the strange week.

We upgraded the model that writes the answers. Read side by side, it was clearly better. Fuller explanations, better handling of multi-step questions, more reliable refusals on trick questions.

And the evaluation score dropped.

What happened when we upgraded only the answering model. The cause sits in the middle of the chain, not at either end: the grader never changed.

We assumed something in the pipeline had broken. One teammate spent a full day checking the document search. Nothing wrong. Finally, we put the old answers, the new answers, and the grader's reasoning next to each other.

The pattern was obvious. The new answers were being marked down for being better. They carried nuance that the grader did not follow, so it read the difference as disagreement. The old model's shorter, flatter answers matched the reference wording almost exactly and scored full marks.

Someone in the room said it best. We had asked a student to grade the teacher's exam.

A student grader behaves in predictable ways.

  • It marks down answers it does not understand

  • It rewards answers that look like what it would have written

  • It cannot tell a confident right answer from a confident wrong one

The result is that scores drift toward the level of the grader. When the answering model becomes more capable than the grader can reliably evaluate, improvements can start looking like regressions in the score. 

The same thing is on the record

MT-Bench, one of the standard papers on LLM evaluation, prints a transcript of its judge getting this wrong.

Question: How many integers satisfy |x + 5| < 10?
Answer A: 19, with no work shown.
Answer B: a careful two-case derivation ending in 20.
Correct answer: 19, the integers from −14 to 4.

The judge was GPT-4, prompted to solve the problem itself before grading. It copied Answer B's derivation almost word for word, miscount included, and picked B. Asked the same question on its own, GPT-4 answers 19 correctly.

The short correct answer lost to the long wrong one. Across 10 math questions, the paper counted judge failures at 14 out of 20 with the default prompt, 6 out of 20 when the judge was told to solve first, and 3 out of 20 when it was given a correct reference answer.

The correct answer was 19, and the judge picked the one that said 20. The plain correct answer lost to the wrong one that showed its working.

We replaced the grader with a stronger model. The new chatbot's scores jumped above the old one, which is what our eyes had been telling us all along.

It kept happening after that too. Below is one of our experiment logs. Look at just one thing: the drop and the recovery have different causes.

Our real experiment log. The fall to 0.41 at runs #36 and #38 came from changes to the chatbot (chunking and output format), but the climb back to 0.71 across #40 to #42 came from fixing the grader, with no change to the chatbot at all. Making it write its reason and stop penalizing formatting differences was enough to move the score on the same answers. Since the grader changed, 0.41 and 0.71 are not directly comparable. The tool has no Japanese UI, so the screen is in English.

This is not just our experience. Engineering teams testing LLM graders report unstable scores across identical samples, gaps in domain knowledge, and answers being overrated simply for containing more surrounding information. Research on judge reliability also finds that position, length, and formatting shape the verdict in ways unrelated to the actual quality of the content.

The rule we took away is simple. The grader must be at least as capable as the thing it grades, and the reference answers themselves must be correct. If you upgrade the chatbot, upgrade the grader first. The grader only runs during testing, so using a stronger model there can be much cheaper than upgrading the production model itself.

What should humans still check?

Even after automating the scoring, we kept one hour a week to read the low-scoring answers together. Tuesday afternoon, usually. It is still the most useful hour of our week.

Here is what that room taught us.

  • Many "wrong" answers were right. A question about opening hours had the reference answer "9 AM to 6 PM." The chatbot said, "Open from 9 in the morning until 6 in the evening, closed on public holidays." The grader called it a mismatch. It was a better answer than our reference. The same fact usually appears in several places in a document set, worded differently, and a good answer can come from any of them.

  • Sometimes our reference answers were wrong. We found three built from an outdated price list. The chatbot was reading the current page and getting punished for being up to date. Now, whenever a client updates their documents, we review the dataset in the same ticket.

  • The grader was fussy about wording. It once docked points because the chatbot wrote "approximately 300 pages" when the book has exactly 300. So we changed the grading prompt to require a written reason before the score. Silly reasons became easy to spot.

Judges also change their minds for reasons that have nothing to do with content

The same paper records a cleaner version of this.

Question: What are some business etiquette norms for doing business in Japan?

Two answers of similar quality were shown to the judge. With the first one placed first, the judge called it better organized and more detailed. Swapping the order flipped the verdict. Nothing about the answers changed.

When the order was swapped, judges reached the same verdict 65.0% of the time for GPT-4, 46.2% for GPT-3.5, and 23.8% for Claude-v1. Claude-v1 picked whichever answer came first in 75% of cases.

The same two answers graded twice, with only the order changed. Nothing about the content changed, and the verdict reversed.

That is why we spend our hour reading the reasons rather than the scores.

What does a drop in the score tell you?

One more story, because it shows what testing buys you.

For months, our chatbot handled text questions well and price questions badly. We assumed price questions were just hard.

Then we added numeric questions to the dataset, like "how many books cost under 2,000 yen," and the failures lined up too neatly. Every failed question involved a table.

The tool that imported web pages had a bug. It silently dropped every table on every page. The chatbot had never seen a single price list. It was not bad at numbers. It was answering with its eyes closed.

You do not find that by typing a few questions and nodding.

What to do this week

Smaller companies are not behind here because of attitude. Japan's Information and Communications White Paper notes that about half of small and medium firms have not set a clear policy on generative AI use at all. The gap is about method, not motivation.

The evaluation cycle we run every week. Note that a human review step survives even after grading is automated, and that the dataset itself gets revised whenever the source documents change.

What we do now works just as well as a buyer's checklist.

  • Keep a list of questions with human-verified answers, and review it whenever the source documents change

  • Always include questions with false assumptions, and questions your documents cannot answer

  • Use a grading model at least as strong as the answering model

  • Require the grader to explain every low score

  • When comparing two answers, grade twice with the order swapped

  • Have people read the low-scoring answers on a regular schedule

  • Judge a change by comparing two versions on the same questions with the same grader, never by one absolute number

  • Reject any change that raises the average score but makes the chatbot start answering out-of-scope questions

So what was actually the hard part?

When we started, we thought building the chatbot was the hard part. It was not. The hard part is knowing whether it is actually good.

Evaluating a chatbot is less like running a benchmark and more like running a school. You need a fair exam with tricky questions, a grader who knows the material, and people who occasionally check the grader's work.

If your company uses a chatbot like this, one question is worth asking the team: "Who grades it, and how do you know the grader is right?" A long silence is also useful information.

If you want to try this on your own documents, run ten questions you already know the answers to through a Kotae trial. (Start here)