Case study · Animarek
Tested 2026-08-04, on a tool whose meeting notes had just shipped.
The test: does it tag who said what correctly, on a real 70-minute call with four people? Not well enough to trust in an automated pipeline. Still better at the words than the tool I am keeping. Usable, if you repair it first. This is the check I now run on any tool before it goes under my business.
The tool may well fix this soon.
I almost adopted it the day I read it.
The transcript was excellent. Full sentences instead of shredded caption fragments. It caught technical terms my usual tool mangles. I run a weekly group call whose transcript feeds a commitments tracker. I was ready to swap it in.
That is how most AI tools get adopted. The output looks good, so the tool must be good. A vendor demo works the same way.
But reading a transcript only tells you if the sentences look right. It cannot tell you if the name in front of each one is right, and the names are the one thing my pipeline depends on.
So I measured it. Two files and some counting. No audio, no lab. The whole test comes down to three moves, and you can run it on any tool you use.
Google Meet knows who spoke, because each person’s mic is a separate recording. It is not listening to a voice and guessing. That is the entire reason it can be the yardstick.
Timestamps drift between tools, so I aligned the two files on words: every run of six consecutive words that appears exactly once in each file. There were 8,039 of them, each with a speaker tag from each tool. The tags either match or they do not. That turns an impression into arithmetic.
Overall tagging accuracy was 84.2%. That sounds survivable, so I cut it into 5-minute windows. Near perfect for 50 minutes, then a collapse, down to 6.4% in the final five. Those minutes are where the group states its weekly commitments. Three of the four commitments landed on the wrong person. An 84% spread evenly is a usable tool. An 84% that is perfect for 50 minutes and then fails is a trap.
With four speakers, random guessing gets the tag right about 25% of the time. This tool scored 6.4%. It cannot be that wrong by accident. The labels had rotated one seat around the table.
Reading could never have caught this. The tool fails fluently.
When Meet gets a word wrong it writes “Buenosides” and you instantly know something broke. When WisprFlow gets a word wrong it writes a clean English sentence nobody said. When it gets a name wrong, you get a believable person saying believable things, under someone else’s name. It genuinely beat my old tool on the words, and that is exactly what makes the failure invisible.
You do not have to trust Meet on this. Some lines give away who said them all by themselves. Two of them:
Meet got both right. WisprFlow got both wrong.
The video editorSo I think my brand... I’m trying to rework my brand on my gaming channel. I have it clear, like, I know what represents me and I know what I like...
MeSo I think my brand... I’m trying to rework my brand on my gaming channel. I have it clear.
MeLike, I know what represents me, and I know what I like...
I gave three AI agents the flawed transcript and a list of who was on the call and what each person owns. Nothing else, no reference file. Their combined fixes took tagging accuracy from 84.2% to 97.0%, and each agent recovered all four weekly commitments on its own. One of them was told nothing about the failure and still found it. So the file is fixable. But the trust is not automatic. Someone has to run the check.
Google Meet stays my transcript of record. It never has to guess.
Most people reading this will never need to check a transcript. That is fine. The habit is the point. A demo is a marketing claim until you have tested the tool in your real conditions: your own data, your own metrics, the thing you actually depend on.
And AI moves fast. This specific flaw may be gone in a few months. Also fine. The numbers expire. The habit of checking does not.
Most of the AI tools going into small businesses right now have been demoed, never measured. I build the measurement, find the specific place a tool breaks, and write down the rule for when it is safe.
Let’s test your stackThis is one call, one group, four non-native English speakers, measured on 2026-08-04. It is a verdict on this tool in my pipeline, not on the tool everywhere. One question stays open: is the cliff about call length, or about the fast cross-talk that starts around minute 50? One short, interruption-heavy call would settle it.