Epure
Founder note

You don't need an observability platform

error tracking vs observability platform

In an incident I open one exception, a count, and one stack. A usage-priced suite bills that spike and gives me a second system to run.

On this pageShow
  1. The outage is the invoice
  2. Two systems
  3. What the count is counting
  4. What this leaves out
  5. When the suite is the right buy
  6. What I run
  7. Questions

You don't need an observability platform.

In an incident I open one exception, a count, and one stack. A usage-priced suite bills that spike and gives me a second system to run.

Checkout started returning 500s. I opened the issue list, not a trace view. The title was TypeError. The count was still moving while I read it. The first in-app frame was chargeCard, our file, our line. I opened the file. The bug was on that line. The rest of the hour was the patch, the rollback check, and a note in the incident channel.

That screen is the whole product I want during an incident. Exception type. A count. One stack. I am not looking for a service map.

An observability platform is a second product sitting on top of that screen. Logs, metrics, traces, session replay, profiling, a retention setting, and a bill. Vendors bundle them because the suite is the business. The bundle is not what I opened. If I cannot get the exception, the count, and our frame from the first screen, the tool has already failed, no matter how many panels it ships.

The outage is the invoice

Usage pricing charges you for events. An outage is a spike of the same event. The hour you most need the tool is the hour the meter runs hottest. I have watched teams pause ingest, or sample harder, in the middle of an incident, because the projected bill moved while they were still reading the stack. The count is the signal. Turning the count down to save money throws away the number that told you the deploy was bad.

A flat self-hosted box does not do this. The disk fills or it does not. You are not negotiating with a meter while chargeCard is throwing. You still operate the box: backups, upgrades, a disk alert. That work is real. It is smaller than paying a surcharge for the fire, and it does not spike because the app did.

Two systems

The app is one system. The suite is another. Someone keeps the suite up: workers, a queue, a config language, an SSO integration, a person who remembers which dashboard is the real one. During the incident you now have two failure modes. The app is throwing, or the watcher is down, or both, and you pick which one to debug first.

I do not want that choice. I want the watcher boring enough that it is not a second on-call. Two containers. A database I already know how to dump. An SDK I already shipped. If the watcher is quiet, I debug the app. If the watcher is loud, I can see why, because there is not a fleet of extra processes hiding the error.

What the count is counting

Grouping is not a clustering model. The default key is a hash of the exception type plus the top in-app frame. The frame part is the file, the function, and the line. Same TypeError from the same line becomes one issue, and the count goes up. Same TypeError from a different function becomes a different issue. You can see why two events did or did not land together. That is the point of a crude key.

It fails in ways you can name.

A minified bundle with no source map makes the in-app frame look like a single letter at line 1. Unrelated throws share a group, or one throw splits across builds because the minified name changed. Fix the map, uploaded for that release. A smarter grouper will not. If the release string on the event does not match the release you uploaded, the stack stays minified and the group key is garbage. Treat that as a release bug.

If two call sites are one bug and the hash splits them, set a fingerprint on the SDK. The server uses a custom fingerprint array when the event sends one, and ignores the default hash. That override is the escape hatch. There is no "similar errors" score. If you need one, you are shopping for a different product.

A spike of one fingerprint can also lie. The count climbs because of a retry loop, not because more users hit the bug. You still want the count. You also want to notice when one issue is eating the pipe. Keep the count, stop storing every copy, and you can see the loop. You do not need a trace for that. You need the same issue, a huge number, and a stack you have already read.

What this leaves out

Session replay shows a cursor on the checkout page. It shows the user clicked Pay. I already knew that. It does not show why chargeCard threw.

A trace waterfall is the right picture when the question is which hop added 200ms and nothing threw. If the request threw, the span wraps an exception I already have. Tracing earns its keep on latency with no error. It does not earn a second system on the day the error is the page.

A CPU profile is a different afternoon. It is not the first screen of a 500.

iOS and Android crash symbol files are their own pipeline. If that is your job, this note is the wrong buying guide. A backend exception store does not symbolicate a mobile crash.

If the SDK still has a traces, replay, or profile sample rate above zero, you are shipping payloads the errors store will not keep. Set those rates to zero when the URL on the other end only stores exceptions. That is a footgun in the SDK config, and it shows up as "we sent it and nothing arrived."

This note is the box I run. If you want someone else to operate the host, that is Cloud, with published prices (Pro $24, Plus $79), and a different decision from the one here.

When the suite is the right buy

Buy the platform when your team already works in traces. Buy it when support watches replays as the normal way they close a ticket. Buy it when the product is mobile and the crash stays unsymbolicated without their pipeline. Buy it when you have a person whose job is the query language and they are good at it. Those are real jobs. They are not the job of "the API is 500ing and I need the stack."

Sentry, the company this SDK comes from, already split the product. Their self-hosted install documents an errors-only compose profile. That profile exists because the rest of the suite is optional. I am not quoting their RAM number here. Their docs do, and if you argue footprint you should read that profile before you repeat a 16 GB figure as if it were the only install. The narrower point: the vendor of the full suite will run errors without the rest, if you ask. Most homepages do not lead with that.

What I run

I run Epure for this job. Two containers, one of them Postgres 16, the app in Rust. The Sentry SDK stays in the app. I change the DSN. It stores exceptions and stacks. It does not store replays, traces, or profiles, because those are not in the product. Pointing a Sentry SDK at a new DSN is envelope ingest, not a claim that the rest of Sentry is on the other end of the URL.

You do not need an observability platform to see a production exception. You need the exception, the count, and one stack with your frame on it. The suite is another product, another bill, and another thing that can page you while the app is already paging you. Keep it if you use it. Drop it when the only thing you opened was the stack.

FAQ

Questions

What do you open first when checkout returns 500?
The issue list. The title is the exception type, the count is still moving, and the first in-app frame names our file and line. The patch starts there. A trace view is a later question, and only if nothing threw.
Why is usage pricing a bad fit for an outage?
An outage is a spike of the same event. The meter runs hottest in the hour you need the count. Teams pause ingest or sample harder while they are still reading the stack. A flat self-hosted box fills disk or it does not. The bill does not move because the app did.
How does Epure group two TypeErrors?
The default key hashes the exception type plus the top in-app frame: file, function, and line. The same type from the same line is one issue. The same type from another function is another issue. A custom fingerprint array on the event replaces that hash. There is no similar-errors score.
When should you buy the full suite?
When the team already works in traces, when support closes tickets from replays, when mobile crashes need their symbolication pipeline, or when someone on the team lives in the query language. Those jobs are real. They are not "the API is 500ing and I need the stack."