
𝗧𝗵𝗲 𝗿𝗲𝗮𝗹 𝗿𝗲𝗮𝘀𝗼𝗻 𝗠𝘂𝗹𝗲𝗦𝗼𝗳𝘁 𝗱𝗲𝘃𝘀 𝗴𝗲𝘁 𝘀𝘁𝘂𝗰𝗸 𝗮𝘁 “𝗷𝘂𝗻𝗶𝗼𝗿 / 𝗺𝗶𝗱” 𝗶𝘀𝗻’𝘁 𝗗𝗮𝘁𝗮𝗪𝗲𝗮𝘃𝗲.
It’s this interview question:
𝗪𝗵𝗮𝘁 𝗵𝗮𝗽𝗽𝗲𝗻𝘀 𝘄𝗵𝗲𝗻 𝘁𝗵𝗲 𝗱𝗼𝘄𝗻𝘀𝘁𝗿𝗲𝗮𝗺 𝗴𝗼𝗲𝘀 𝗱𝗼𝘄𝗻… 𝗮𝗻𝗱 𝘆𝗼𝘂𝗿 𝗾𝘂𝗲𝘂𝗲 𝗸𝗲𝗲𝗽𝘀 𝗳𝗶𝗹𝗹𝗶𝗻𝗴?
Most answers sound like this:
“Use async + Anypoint MQ.”
That’s a good start.
But it’s not protection.
Because async doesn’t control pressure. It delays it.
𝗛𝗲𝗿𝗲’𝘀 𝘁𝗵𝗲 𝗺𝗮𝘁𝗵 𝘁𝗵𝗮𝘁 𝗰𝗵𝗮𝗻𝗴𝗲𝘀 𝘁𝗵𝗲 𝗰𝗼𝗻𝘃𝗲𝗿𝘀𝗮𝘁𝗶𝗼𝗻:
10,000 msgs/day ≈ 416 msgs/hour
8 hours downtime → ~3,333 messages backlog
Now the “recovery” begins:
👾 vCore spikes while the consumer tries to catch up
👾 retry storms + log noise
👾 the downstream gets hammered right when it’s most fragile
👾 and you end up doing manual cleanup at 3 AM
𝗦𝗲𝗻𝗶𝗼𝗿 𝗮𝗻𝘀𝘄𝗲𝗿 (𝗮𝗻𝗱 𝗿𝗲𝗮𝗹-𝘄𝗼𝗿𝗹𝗱 𝗱𝗲𝘀𝗶𝗴𝗻):
✅ 𝗖𝗶𝗿𝗰𝘂𝗶𝘁 𝗕𝗿𝗲𝗮𝗸𝗲𝗿
Pause consumption the moment downstream health drops.
✅ 𝗣𝗿𝗶𝗼𝗿𝗶𝘁𝘆 𝗗𝗟𝗤
Failures go aside for controlled replay (P1 first).
✅ 𝗜𝗱𝗲𝗺𝗽𝗼𝘁𝗲𝗻𝗰𝘆 (𝗯𝗼𝗻𝘂𝘀)
So replay doesn’t create duplicates.
That’s it.
You’re not “building integrations.”
You’re building controlled failure.
I put the high-level architecture + explanation into a short PDF (based on what we use in real systems).