Use cases

Sort log lines and alerts

Which service, how bad, and whether customers feel it, on every alert, so on-call sees the one that matters.

Try it on this example

Example · A payments alert: every checkout failing on a ledger timeout

Log line or alert text

2026-09-14T19:42:07Z [FIRING] payments-api ERROR ledger-svc POST /v2/commit timeout after 30000ms (retry 3/3 exhausted) order=ORD-88213 request_id=7f3c9a1e-2b44 merchant=m_20417 | alert: checkout_commit_failures 412/412 customer checkouts failed in last 5m
  1. Which of our services is failing or affected in this log line?Payments95%
  2. How severe is the problem this log line describes?Outage100%
  3. Does the problem in this log line reach customers or merchants?Yes98%
  4. What kind of failure does this log line show?Timeout or no response100%
  5. Does the log line say the problem has recovered or a retry succeeded?No95%
  6. Does the log line contain a password, key, token or a customer's personal data?No85%

These are real answers stored from one run on this example.

The prism behind it

Sort log lines and alerts6 questions

Fields

  • Log line or alert text

Context

The lines come from the error logs and alert history of Parcelry, an online shop platform, exported for the weekly on-call review. Each line is one log entry or one alert message as the monitoring tool wrote it. The answers sort a week of lines by service and severity, so the review starts with the ones that hurt customers, and flag lines that leak secrets or personal data. Rates, counts across lines and whether an alert is new are worked out in code from the alert history. Our services: - Payments: payments-api, ledger-svc, payouts-worker, the card provider connection. - Auth: auth-svc, login, SSO, session and token handling. - Search: search-api, indexer, product search. - Orders: orders-api, checkout-ui, cart, stock reservations. - Notifications: email, SMS and push senders, webhooks to merchants. - Platform: databases, queues, caches, Kubernetes nodes, load balancers, DNS and certificates.

Questions

  1. Which of our services is failing or affected in this log line? Choice

    Use the service list in the context. When one service reports a failure in another that it calls, pick the service that failed. Pick one option.

    • Payments Taking, recording or paying out money, including the ledger.
    • Auth Logging in, sessions, tokens and single sign-on.
    • Search Product search and the search index.
    • Orders Checkout pages, carts, orders and stock reservations.
    • Notifications Emails, texts, push messages and webhooks sent to merchants.
    • Platform Shared infrastructure that no single service owns: databases, queues, caches, nodes, load balancers, DNS and certificates.
  2. How severe is the problem this log line describes? Scale

    Go by what the line says happened, not by its log level alone: an ERROR line about one retried request can be a warning, and an INFO line can report an outage.

    • Info A normal event or a routine message, with nothing failing.
    • Warning Something is degraded or near a limit, or a failure was retried and succeeded, but requests are still served.
    • Error Some requests or jobs fail, while the service still works for most.
    • Outage A service or function is down for all or most requests, or the line says all attempts failed.
  3. Does the problem in this log line reach customers or merchants? Yes / No

    Count failed checkouts, logins, searches, payments, orders or messages that a shopper or merchant would see. A failure in a background job with no sign it reached anyone does not count. Yes: The line shows customers or merchants are affected. No: The line shows no effect on customers or merchants.

  4. What kind of failure does this log line show? Choice

    Go by what the line says went wrong. When a timeout comes from a dependency, it is still a timeout. Pick one option.

    • Timeout or no response A call or job took too long, or a dependency did not answer.
    • Error response A call returned an error status or an exception, other than a timeout.
    • Resource limit Memory, disk, CPU, connections, a quota or a rate limit ran out or came close.
    • Crash or restart A process, pod or node crashed, was killed or restarted.
    • Configuration or deploy A bad setting, a missing secret, an expired certificate or a failed deploy or migration.
    • Security event Failed logins in bulk, blocked requests, permission denied or a suspected attack.
    • No failure The line reports a normal event or a recovery, with nothing failing.
  5. Does the log line say the problem has recovered or a retry succeeded? Yes / No

    Count words such as resolved, recovered, back to normal, healthy again or retry succeeded. Yes: The line says the problem has cleared. No: The line says the problem is ongoing, or does not say.

  6. Does the log line contain a password, key, token or a customer's personal data? Yes / No

    Count a password, an API key, a bearer or session token, a card number, or a name, email address, phone number or postal address of a customer. Order numbers, request IDs and masked values such as "tok_****" do not count. Yes: The line contains at least one such value. No: The line contains no such value.

Lens columns

service, service_probability, severity, severity_average, customer_facing, customer_facing_probability, failure_type, failure_type_probability, recovered, recovered_probability, leaks_secret_or_personal_data, leaks_secret_or_personal_data_probability

Run it on your own text

Add this prism in the app, change any question, and test it on a file of your own.

Ask for an invite