Near-Real-Time Activity System with Infinite Scroll
An activity feed that got slower the more you used it. Power users were being punished for showing up.
Senior Software Engineer
CV ↗An activity feed that got slower the more you used it. Power users were being punished for showing up.
Moving money between four account types sounds simple. Some of them won’t even talk to each other directly.
Greenfield build: users needed to move funds across four account types (external bank, spending wallet, savings wallet, and a trading account), where not every pair connects directly. The system had to feel close to instant, support both a “give me the result now” and “handle this in the background” mode, auto-recover from failed steps at the Green Dot layer, fail gracefully when a step can’t recover, reconcile against Green Dot end of day, and page someone the moment anything failed in a way that mattered.
A new bank-transfer-service: one Lambda run as a service, where the event’s action picks the job, paired with a single queue, BankTransferServiceQ. The app reaches it through API Gateway; internal services invoke it directly.
The core model split every transfer into two entities: an Instruction (1:1 with the user’s request) and one or more Transactions (n:1 to the instruction, one per hop). Everything downstream operates on this split, regardless of whether the transfer runs sync or async.
The processing chain:
createTransfer: the only way in. With async: false, it chains createInstruction → (processInstruction → processTransaction → updateTransaction) in a loop until the instruction is COMPLETE or has to wait on a webhook. With async: true, it returns as soon as the instruction exists and the queue takes it from there. Either way the caller gets a 200 with the instruction’s status.createInstruction: saves the instruction as SUBMITTED. In async mode, every hop from here on goes through the queue.processInstruction: determines the required hop sequence for this transfer, persists each hop’s state as INITIAL, creates the first transaction record as SUBMITTED, and updates the instruction’s hop state in a single TransactWrite. Triggers processTransaction next. Once every hop is complete, marks the instruction COMPLETE and exits.processTransaction: submits the transaction to Green Dot, which also handles the trading account’s side with Apex. If Green Dot completes it instantly, moves straight to updateTransaction.updateTransaction: sets the transaction to its terminal state, then re-triggers processInstruction for the next hop. It fires either straight from processTransaction, or when the settlement webhook lands: bank-events drops it onto the same BankTransferServiceQ, and onEventReceived calls updateTransaction.Internal, system-initiated transfers (like recurring investment debits) invoke createTransfer directly, with async: true where nobody needs to wait.
Every function pairs with its own healing job on an EventBridge rate() schedule, putting anything stuck in a non-terminal state back on the queue with exponential backoff. An end-of-day job checks against Green Dot for missed updates and status mismatches, and fixes what it finds.
If a hop still fails after 5 retries, its transaction goes VOID and the instruction FAILED. If an earlier hop has already moved money, nothing rolls back on its own: the instruction lands in SUPPORT_REQUIRED, and support can re-trigger the hop once Green Dot is back. That takes Green Dot being down for a day, or bad code. Logs carry specific keys, a Datadog monitor watches for them, and PagerDuty pages someone once a threshold is crossed, including for anything in SUPPORT_REQUIRED.
A fully async, multi-hop system is genuinely harder to observe than a linear one: there’s no single call stack to trace. The trade-off was accepted deliberately, and covered by disciplined per-step logging so every hop’s state is traceable even though the execution path itself isn’t linear.
One codebase, one set of logic, handling both real-time and background transfers identically underneath. End-to-end transfer completes in under 5 seconds, excluding actual ACH settlement time. Immutable transaction records, one-directional state transitions, and active PagerDuty monitoring for anything that fails past the point self-healing can fix.
Card swipes, approvals and settlements arrive whenever they like. Someone’s balance can’t be “eventually correct”.