Case study · Animarek

I tested if the new meeting note taker from WisprFlow is to be trusted in my pipeline. It isn’t.

Tested 2026-08-04, on a tool whose meeting notes had just shipped.

The test: does it tag who said what correctly, on a real 70-minute call with four people? Not well enough to trust in an automated pipeline. Still better at the words than the tool I am keeping. Usable, if you repair it first. This is the check I now run on any tool before it goes under my business.

84.2%
tagging accuracy across the whole call
6.4%
in the final five minutes
97.0%
after a repair pass with no reference file

The tool may well fix this soon.

100%75502500-5 min · 86.4% tagging accuracy · 457 anchors5-10 min · 99.3% tagging accuracy · 696 anchors10-15 min · 100.0% tagging accuracy · 546 anchors15-20 min · 100.0% tagging accuracy · 817 anchors20-25 min · 95.7% tagging accuracy · 704 anchors25-30 min · 98.6% tagging accuracy · 567 anchors30-35 min · 100.0% tagging accuracy · 828 anchors35-40 min · 91.0% tagging accuracy · 523 anchors40-45 min · 95.0% tagging accuracy · 525 anchors45-50 min · 96.2% tagging accuracy · 650 anchors50-55 min · 74.2% tagging accuracy · 462 anchors55-60 min · 51.0% tagging accuracy · 492 anchors60-65 min · 9.7% tagging accuracy · 442 anchors65-70 min · 6.4% tagging accuracy · 330 anchorsthe break · ~50 min86.4%100%74.2%51.0%9.7%6.4%010203040506070 min
above the random-guess linebelow it25%: random guess, four speakers
Tagging accuracy in 5-minute windows. Near perfect for 50 minutes, then a cliff. The dashed line is what random guessing scores with four speakers. The last two windows sit below it.

I almost adopted it the day I read it.

The transcript was excellent. Full sentences instead of shredded caption fragments. It caught technical terms my usual tool mangles. I run a weekly group call whose transcript feeds a commitments tracker. I was ready to swap it in.

That is how most AI tools get adopted. The output looks good, so the tool must be good. A vendor demo works the same way.

But reading a transcript only tells you if the sentences look right. It cannot tell you if the name in front of each one is right, and the names are the one thing my pipeline depends on.

So I measured it. Two files and some counting. No audio, no lab. The whole test comes down to three moves, and you can run it on any tool you use.

The test, in three moves

1

Compare against something that does not guess.

Google Meet knows who spoke, because each person’s mic is a separate recording. It is not listening to a voice and guessing. That is the entire reason it can be the yardstick.

2

Count, do not eyeball.

Timestamps drift between tools, so I aligned the two files on words: every run of six consecutive words that appears exactly once in each file. There were 8,039 of them, each with a speaker tag from each tool. The tags either match or they do not. That turns an impression into arithmetic.

3

Slice it, because an average hides a cliff.

Overall tagging accuracy was 84.2%. That sounds survivable, so I cut it into 5-minute windows. Near perfect for 50 minutes, then a collapse, down to 6.4% in the final five. Those minutes are where the group states its weekly commitments. Three of the four commitments landed on the wrong person. An 84% spread evenly is a usable tool. An 84% that is perfect for 50 minutes and then fails is a trap.

Too wrong to be random

With four speakers, random guessing gets the tag right about 25% of the time. This tool scored 6.4%. It cannot be that wrong by accident. The labels had rotated one seat around the table.

100%95%58%42%MeThe video editorThe designerThe accountant✓ 100%
The final ten minutes. Arrows show whose name the tool put on each person’s words. The video editor became me, the designer became the video editor, my words split between two others. Only the accountant kept their own label.

Reading could never have caught this. The tool fails fluently.

When Meet gets a word wrong it writes “Buenosides” and you instantly know something broke. When WisprFlow gets a word wrong it writes a clean English sentence nobody said. When it gets a name wrong, you get a believable person saying believable things, under someone else’s name. It genuinely beat my old tool on the words, and that is exactly what makes the failure invisible.

You do not have to trust Meet on this. Some lines give away who said them all by themselves. Two of them:

  • “...what [Mati] put... is the difference.” Someone is talking about me. The tool says I said it. Nobody talks about themselves in the third person.
  • “One thing that came to mind when [the video editor] was speaking...” The tool tagged this line as the video editor. Same mistake: you do not talk about yourself in the third person.

Meet got both right. WisprFlow got both wrong.

Google Meetright speaker

The video editorSo I think my brand... I’m trying to rework my brand on my gaming channel. I have it clear, like, I know what represents me and I know what I like...

WisprFlowwrong speaker

MeSo I think my brand... I’m trying to rework my brand on my gaming channel. I have it clear.

MeLike, I know what represents me, and I know what I like...

WisprFlow put this in my mouth. I don’t have a gaming channel. The video editor does. Both panels trimmed equally for length.

Repairable, which changes the verdict

I gave three AI agents the flawed transcript and a list of who was on the call and what each person owns. Nothing else, no reference file. Their combined fixes took tagging accuracy from 84.2% to 97.0%, and each agent recovered all four weekly commitments on its own. One of them was told nothing about the failure and still found it. So the file is fixable. But the trust is not automatic. Someone has to run the check.

Safe as-is

  • Solo recordings and one-on-ones.
  • Anything I read myself, where I already know who spoke.

Not safe without a repair pass

  • Anything feeding a commitments log or a person’s record.
  • Group calls past about 45 minutes.
  • Groups where several speakers have similar accents.

Google Meet stays my transcript of record. It never has to guess.

Trust, but verify

Most people reading this will never need to check a transcript. That is fine. The habit is the point. A demo is a marketing claim until you have tested the tool in your real conditions: your own data, your own metrics, the thing you actually depend on.

And AI moves fast. This specific flaw may be gone in a few months. Also fine. The numbers expire. The habit of checking does not.

Most of the AI tools going into small businesses right now have been demoed, never measured. I build the measurement, find the specific place a tool breaks, and write down the rule for when it is safe.

Let’s test your stack

This is one call, one group, four non-native English speakers, measured on 2026-08-04. It is a verdict on this tool in my pipeline, not on the tool everywhere. One question stays open: is the cliff about call length, or about the fast cross-talk that starts around minute 50? One short, interruption-heavy call would settle it.