Near-Real-Time Activity System with Infinite Scroll
An activity feed that got slower the more you used it. Power users were being punished for showing up.
Senior Software Engineer
CV ↗An activity feed that got slower the more you used it. Power users were being punished for showing up.
Moving money between four account types sounds simple. Some of them won’t even talk to each other directly.
Card swipes, approvals and settlements arrive whenever they like. Someone’s balance can’t be “eventually correct”.
Not all Green Dot interactions are synchronous. Account creation, transfer completion, and card-street transaction notifications all arrive as webhooks, on Green Dot’s timeline, not ours. In a banking product, this creates a real correctness risk: a card transaction hitting a balance can race an in-app transfer, and if events aren’t processed near-instantly and in the right order, users see incorrect balances, even briefly. This system also owns triggering the next hop in the multi-step transfer flow.
An incoming webhook lands on an ECS-hosted saveEvent endpoint, gets normalized into our internal shape, and is saved to a BankEvent DynamoDB table with state SAVED.
That write triggers a DynamoDB Stream, picked up by a consumeEvent Lambda, which flips the state to PROCESSING and drops the event onto a BankEvents SNS topic.
The topic is filtered by event type (USER, ACCOUNT, TRANSFER, CARD, and so on), fanning out to a dedicated SQS queue per domain (BankUserServiceQ, BankAccountServiceQ, BankTransferServiceQ, BankCardServiceQ). Each queue runs a 2-minute visibility timeout with a max receive count of 5 before giving up and escalating.
Each service owns its own onEventReceived handler, does whatever domain-specific work the event requires, and marks the BankEvent COMPLETE on success.
If a queue exhausts its retries, the message lands in that service’s own dead-letter queue (e.g. BankTransferServiceDLQ), where a separate handler marks the event FAILED and logs it for threshold-based alerting.
FAILED event for support or engineering to investigate directly, rather than relying on someone noticing a symptom downstreamSame trade-off as the transfer system: a fully async, fan-out architecture is harder to trace end-to-end than a linear call stack. Mitigated the same way: disciplined state tracking (SAVED → PROCESSING → COMPLETE/FAILED) at every stage, so any event’s status is queryable directly instead of inferred.
A near-real-time event loop that recovers from its own failures, retrying automatically where possible, surfacing what it can’t fix within the hour instead of letting it sit silent.