Skip to content
Bernhard Götzendorfer
Debugging & Solutions

46 Seconds Inside a 30-Second Webhook

A render call could take 46 seconds inside a webhook with a 30 second budget. The second finding turned out to be worse than the first one.

TL;DR

In EventDrop you can buy a recap, an automatically edited video made from the photos of an event. The render was triggered directly inside the Stripe webhook, awaited, in the middle of the request path. Stripe reports the payment through a webhook, an automatic callback into my server that has to answer inside a tight time budget. Counted up from the source, that call takes roughly 46 seconds in the worst case. The time budget of the webhook route is 30. That was the first finding. The second one was worse: a red health check immediately wrote the status failed, and the nightly recovery only picks up rows with status paid. A two-minute renderer restart therefore killed every recap paid for in that window, permanently.

What a Recap Is and Why the Webhook Triggered It

A recap is a paid extra: the photos of an event are automatically cut into a short video, with music, in an order that makes reasonable sense. The rendering itself does not run in the web application but in a separate service, which the web application only kicks off.

The obvious place to do that is the moment of payment. Stripe reports via webhook that the payment went through, and that is exactly where the call sat that starts the render. Immediacy was the intention: the customer pays, the video starts coming into existence. And it looked fine for a long time, because in the normal case this dispatch takes milliseconds. A POST to the render service, a confirmation back, done.

The problem was never the normal case. It was the question of what happens when the render service does not answer right away, and I had never calculated that question, only estimated it. I describe my tools and the way I work in the tooling section, but the best set of tools does not replace the one calculation you never did.

The Attempt

The function in question is called triggerRecapRender. It had three awaiting callers:

  1. the Stripe webhook, when a recap was paid for,
  2. a server action used to redeem a free recap,
  3. the nightly recovery, phase 4b, which catches up on stranded cases.

Of those three, exactly one had a time budget that fit the job: the nightly recovery is allowed to run for 300 seconds. That is the whole shape of the defect. The same function, called three times, but only one of those places had room for its behaviour.

When I built it, I was thinking in responsibilities rather than time budgets: payment happens here, so the start happens here. A clean story and a poor design, because it puts a slow foreign operation into a path that has to be fast.

The Wall

The calculation is fully present in the source, you only have to add it up once:

health precheck          5 s   (HEALTH_TIMEOUT_MS = 5_000)
attempts 1 through 4  4 x 10 s (RENDER_TIMEOUT_MS = 10_000)
backoff in between   2 + 4 + 8 s (baseDelay 2_000, maxDelay 8_000, jitter up to +50%)
------------------------------------------------
worst case             roughly 46 s

Against that stands one line in the deployment configuration: the Stripe webhook route has maxDuration 30.

46 against 30. That is not close. In the decision document I later put it like this: "This makes the worst case in the 30 s webhook not tight but structurally impossible to meet. A timeout there is not a cosmetic delay: the webhook answers with 5xx or not at all, and Stripe redelivers for three days."

Those three days turn a performance topic into an operations topic. A webhook that runs into a timeout is a failed delivery attempt from the payment provider's point of view. It comes back. And back again, against an endpoint that returns the same result every time. The fault was entirely mine: 30 seconds is a sensible budget for a webhook. I had simply put something in there that does not belong.

The Diagnosis

The second finding came from reading the same code, and it is the actual reason for this article.

When the health precheck was red, meaning the render service was not answering, the function immediately wrote the status failed into the recap row. Not pending, not retry, but failed. That sounds sensible until you know the other half: the nightly recovery, phase 4b, only looks for rows with status paid. A row sitting on failed never appears in that query again.

So failed was a one-way street. The decision document puts it like this: "A row cleared away like that is therefore never caught up, the customer has paid, the video does not exist, and the only way back is a manual admin retry. A two-minute restart of the render service therefore cost every recap paid for in that window."

Two minutes. A deployment of the render service takes longer than that. Every ordinary restart, every update, every brief unavailability would have written everything paid for in that window permanently into the terminal state failed, with no alert and nobody noticing except the customer whose video never arrived.

That is exactly the distinction this is about. The timeout was the visible problem. The quiet finality was the expensive one. I found neither through an outage but by recalculating against the source, the same exercise as back when I let agents audit my own website: measure instead of assume.

The first finding cost seconds. The second one turned a two-minute restart into a permanent loss for everyone who had paid in that window.

The Fix: The Trigger Leaves the Request Path

The answer is old and unspectacular: an outbox. The request now does exactly one thing, which is to write a row. A worker running every two minutes picks it up and performs the dispatch, with a time budget that fits.

The obvious variant would have been to reuse the existing outbox for upload side effects. That failed in two places. Its upload_id column is NOT NULL, but a recap has no upload as its parent. And worse: the processing path there treats a row with no findable upload as done. "A recap row with upload_id IS NULL falls straight into exactly that branch, the dispatch would be a silent no-op, reported as success." A nullable foreign key would not have fixed the defect, it would have hidden it.

What it became is a sibling table, recap_side_effects, with the same shape as the existing one and a parent of its own:

CREATE TABLE recap_side_effects (
  recap_id      uuid NOT NULL REFERENCES recap_videos(id) ON DELETE CASCADE,
  effect_type   text NOT NULL,
  status        text NOT NULL DEFAULT 'pending',
  attempts      int  NOT NULL DEFAULT 0,
  next_attempt_at timestamptz NOT NULL DEFAULT now(),
  UNIQUE (recap_id, effect_type)
);

The UNIQUE key makes enqueuing idempotent, built so that a second attempt does nothing twice: a second webhook for the same payment does not create a second row. Claiming runs through a database function with FOR UPDATE SKIP LOCKED, so two parallel workers cannot grab the same row. And next_attempt_at defaults to now, because here the outbox is the only path and not a straggler next to a direct call.

The second part of the fix is an error contract. Before, there were only errors. Now there are transient and permanent ones. A red health check, a network error, a timeout and a 5xx from the render service are transient and lead to another attempt. A 4xx, a missing configuration, an event that cannot be found, too few photos are permanent. Only the dead path after five attempts still writes failed.

The default there is deliberately the conservative one: "Everything that is not explicitly classified as transient stays permanent, the default is therefore today's behaviour, and a new error source behaves exactly as before until somebody classifies it." A new error class should not behave worse than the state before the rebuild just because nobody has sorted it yet.

The nightly phase 4b did not disappear. It was rebuilt: instead of triggering itself, it now adds a missing outbox row or wakes a stuck one. It is the backstop for the one hole that stays open.

What the Outbox Pattern Solves in General

The trick is a separation that sounds trivial in hindsight. "The job has been accepted" and "the job has been done" are two different promises, and the request only has to guarantee the first. That turns retry, backoff and giving up into properties of data instead of control flow inside a handler whose time budget somebody else sets.

The second gain is visibility. A stuck side effect used to be a state inside a finished process, which is to say nothing at all. Now it is a row with an attempt counter and a next date, one you can look at and re-arm. That is related to what interested me about DevWatchdog: states nobody touches any more are harmless only because they are invisible.

The honest costs sit next to it: up to two minutes of waiting, a small window between the payment and the row, and corpses in status processing when a worker dies mid-run. For that there is a 15-minute threshold and a re-arm path: "The window is therefore not zero, but bounded and self-healing, before it was unbounded and never healed."

And a limit belongs with it. This is not an ideology about message queues. It answers exactly one question, namely who carries the responsibility for a slow side effect. Where the side effect is fast and reliable, the direct call remains the better option.

What I Am Taking Away

  1. A time budget is a contract, not a guideline. If you hang a retry loop inside a function with a 30-second budget, you have to calculate the loop, not estimate it. Five seconds plus four times ten plus fourteen seconds of waiting is a calculation that takes two minutes, and I never did it.

  2. The expensive question is not what happens on failure. The expensive question is what happens when that failure state is final. Here failed was a one-way street, and the only recovery that existed was looking the other way.

  3. An error needs a class before it gets a reaction. Transient and permanent demand opposite behaviour. And the default for everything not yet classified has to be the old, known behaviour, otherwise the rebuild quietly makes something worse that used to work.

  4. Two things with different parents belong in two tables. The nullable foreign key would not have saved a row here, it would have turned the dispatch into a silent success. A second table with the same shape is cheaper than a shared one with an exception.

  5. A green pipeline is no evidence that the migration ran. That is why the deploy was migration-first: the workflow, then a look at the ledger, then the push. The same stance as in Verification, Not Typing, only with a database instead of an editor.

Conclusion

The rebuild covered 24 files, 2,436 inserted lines, a migration of 178 lines and 33 tests that were red against the old state. That is a lot for one call hung in the wrong place. But it was not the call that was expensive, it was the state behind it.

And the limits belong here, otherwise this would be a success report rather than a write-up. The migration could only be checked statically, because Docker was not running on my machine that evening. The worker deliberately processes one recap per run, because a single dispatch costs up to 46 seconds against a 60-second budget. And the enqueue is still not transactional. Open points, not solved ones.

What I ask first has changed. No longer "what happens if this fails" but "what happens when this state is final, and who ever looks at it again". The road from a weekend build to something that has to run every day consists pretty much of questions like that, as I have described elsewhere.