Skip to content
Blog

Why your error tracker keeps grouping unrelated bugs into one issue

Your error tracker's issue list shows one item with a rising count, but the events inside it are not the same bug. Here is why grouping goes wrong, how the major tools each explain it, and what an honest fix looks like: normalize the message, drop line numbers, hash in-app frames, and give the caller an escape hatch.

You open the issue list and one item has 4,000 events and a title like TypeError: Cannot read properties of undefined. You click in expecting one bug. Instead the stack traces underneath are three or four genuinely different failures that happen to throw the same exception type from nearby code. Somebody fixes the loudest one, the count keeps climbing, and the tracker looks broken even though it did exactly what it was built to do.

This is not a bug in your error tracker. It is what "grouping" means, applied to code it has never seen, and it fails in the same handful of ways in every tool that does it. This post is about why, using Sentry's own public documentation as the worked example (it is the tool most people hit this with first), and what a fingerprinting algorithm has to do to make the failure rarer.

What "grouping" is actually doing

No error tracker stores one row per bug. It stores one row per event (an occurrence) and decides, at ingest, which issue that event belongs to. The decision is a fingerprint: a string computed from the event, hashed, and matched against existing issues. Same fingerprint, same issue. Different fingerprint, a new issue is created.

Sentry's own documentation on event grouping describes the default order plainly: check for a custom fingerprint first, then fall back to "stack trace, exception type and value," and finally the message if there is no usable stack. Every other tracker that groups events does some version of the same thing, because the alternative, one row per event with no grouping at all, makes an issue list useless within a day.

The failure mode follows directly from what the fingerprint is built from. Two things cause almost every case of "these are different bugs but they're one issue":

  • The call site is the same, the cause is not. A shared utility function called from a dozen places throws the same exception type with a similar message no matter which caller passed it bad data. If grouping weights the throw site over the caller, everything routed through that utility becomes one issue regardless of why it failed.
  • The message looks identical after normalization. Grouping deliberately treats "user 812 not found" and "user 4409 not found" as the same shape, because otherwise every distinct id would fork a new issue and the count would be meaningless. That normalization is correct for the case it targets and wrong the moment two different failures happen to produce structurally similar messages with different meanings.

Both are the tracker doing its job. The problem is that "the job" is a judgment call about what "the same bug" means, made by an algorithm with no idea what your code is supposed to do.

The fix every serious tracker converges on

Sentry's documentation lists the escalating set of tools it gives you once the default gets it wrong: merge issues that were wrongly split, write a fingerprint rule that reassigns incoming events based on a matcher, write a stack trace rule, or set a fingerprint directly in the SDK call that reports the error. That last one, a custom fingerprint passed by your own code, is the actual fix for the shared-utility case: you know that a payment failure and a signup failure are different bugs even though they both throw from the same validation helper, so you tell the tracker that directly instead of hoping the algorithm infers it.

Every tracker that does automatic grouping ships some version of this same escape hatch, because the alternative is either grouping too coarsely (this problem) or too finely (a one-line code change starts a new issue for what is the same bug, which is its own kind of noise). There is no algorithm that gets both right on code it has never seen. The honest version of the product tells you the rule it used and lets you override it; the dishonest version hides the rule and asks you to trust the count.

What this looks like, concretely, in RealUptime Errors

We can describe our own fingerprinting exactly, because it is a few hundred lines of deterministic code, not a proprietary black box, and it groups on three things:

  1. The exception type, verbatim. TypeError and RangeError never group together.
  2. The message, normalized. Quoted strings become "?", UUIDs and long hex runs become #, and any remaining digit run becomes #. user 812 not found and user 4409 not found hash to the same normalized message; key "abc-123" is stale and key "def-456" is stale do too. This is the same trade Sentry describes: it collapses variable data on purpose, and it is the exact mechanism that can over-group two different failures that both happen to say "user # not found".
  3. The in-app stack frames, by file path and function name, with line numbers deliberately dropped. A one-line edit above the throw site should not start a new issue; that is also why a shared utility called from several places can pull unrelated failures into one issue when the top in-app frame is the utility itself rather than the caller.

The escape hatch is the same one Sentry documents: call the SDK with an explicit fingerprint: [...] array, and that value is hashed as given, bypassing the message and stack logic entirely. If a shared validator throws for five different reasons, pass a fingerprint that includes the actual failure reason and the five reasons become five issues instead of one.

When grouping has already gone wrong on history that already shipped, an issue merge repairs it after the fact: existing events move from the wrong issue to the right one in one transaction, and future events with the old fingerprint are redirected automatically, so the fix does not require replaying data. This matters because a fingerprint change alone (a new custom fingerprint, a fix to the normalization) only affects events from that point forward; without a merge tool, the history stays split or stays wrongly combined.

The fingerprint algorithm also carries a version number on every issue it creates. If the algorithm itself changes later, that is a new version, not a silent regrouping of history you already triaged. That is a smaller thing than it sounds, and it is also the difference between "the count changed because of a real event" and "the count changed because we reshuffled the math," which is exactly the kind of quiet drift the honesty register we hold ourselves to does not allow.

How to tell if this is happening to you right now

Open your loudest issue and read five events from it, not one. Specifically:

  • Do the stack traces share the exact same top in-app frame, or just the same exception type?
  • Does the message differ only in a number, id, or quoted value, or does it differ in the actual noun ("user not found" vs. "session not found" both normalizing to the same shape)?
  • If you have deploy or release data attached, did events with meaningfully different causes start on different releases but land in the same issue?

If the answer to any of those is yes, you have found an over-grouped issue, and the fix is a custom fingerprint on the call sites that need to be told apart, not a bigger investigation into "why is the count so high."

The honest limit

No fingerprinting scheme, ours included, can read intent. It can only read the shape of a message and a stack trace, and shape is a proxy for cause, not cause itself. The tools that help (custom fingerprints, merge, versioned algorithms so nothing regroups history silently) exist because the proxy is known to be imperfect, not because anyone found a way to make it perfect. A tracker that claims its default grouping is always right is the one to be skeptical of; the honest claim is narrower: the default gets most cases right, tells you what it used, and gives you a way to correct the rest.

If you are choosing or evaluating an error tracker and this is the question that matters to you, RealUptime Errors groups on exactly the algorithm described above, with a free tier of 10,000 events a month to try it against your own noisy code before paying for anything. See how the free tier compares to Sentry's, or check the pricing page for what the paid tiers add on top.

Sources: Sentry, event grouping, read September 21, 2026.