Amol Sood

Senior Software Engineer

CV ↗
How deep do you want to go?
02

Real-Time, Self-Healing Transfer System (Green Dot integration)

Moving money between four account types sounds simple. Some of them won’t even talk to each other directly.

The problem

Greenfield build: users needed to move funds across four account types (external bank, spending wallet, savings wallet, and a trading account), where not every pair connects directly. The system had to feel close to instant, support both a “give me the result now” and “handle this in the background” mode, auto-recover from failed steps at the Green Dot layer, fail gracefully when a step can’t recover, reconcile against Green Dot end of day, and page someone the moment anything failed in a way that mattered.

Constraints

  • ACH settlement isn’t instant (deposits take 3 business days, withdrawals take 1), and a webhook signals completion, which is what triggers the next hop
  • Money can only enter or leave the spending/savings pair through the spending wallet specifically; that single rule is what determines whether a transfer needs 1 or 2 hops
  • All logging had to be redacted: no raw sensitive transfer data in logs
How it flowsHover or tap any piece to see why it's there
Mobile appusersAPI GatewayHTTPS entryRecurring investmentsinternal servicecreateTransferbank-transfer-serviceKeyed logsevery actionDatadog monitorthresholdPagerDutyon callcreateInstructionbank-transfer-serviceBankTransferInstructionDynamoDB tableBankTransferServiceQSQSprocessInstructionbank-transfer-serviceBankTransferTransactionDynamoDB tableprocessTransactionbank-transfer-serviceGreen Dotwallets, bank, Apexhealing jobsbank-transfer-serviceupdateTransactionbank-transfer-serviceonEventReceivedbank-transfer-servicebank-eventsTRANSFER webhooksrate() schedulesEventBridge Schedulerend-of-day checkbank-transfer-service

What I built

A new bank-transfer-service: one Lambda run as a service, where the event’s action picks the job, paired with a single queue, BankTransferServiceQ. The app reaches it through API Gateway; internal services invoke it directly.

The core model split every transfer into two entities: an Instruction (1:1 with the user’s request) and one or more Transactions (n:1 to the instruction, one per hop). Everything downstream operates on this split, regardless of whether the transfer runs sync or async.

The processing chain:

  • createTransfer: the only way in. With async: false, it chains createInstruction → (processInstruction → processTransaction → updateTransaction) in a loop until the instruction is COMPLETE or has to wait on a webhook. With async: true, it returns as soon as the instruction exists and the queue takes it from there. Either way the caller gets a 200 with the instruction’s status.
  • createInstruction: saves the instruction as SUBMITTED. In async mode, every hop from here on goes through the queue.
  • processInstruction: determines the required hop sequence for this transfer, persists each hop’s state as INITIAL, creates the first transaction record as SUBMITTED, and updates the instruction’s hop state in a single TransactWrite. Triggers processTransaction next. Once every hop is complete, marks the instruction COMPLETE and exits.
  • processTransaction: submits the transaction to Green Dot, which also handles the trading account’s side with Apex. If Green Dot completes it instantly, moves straight to updateTransaction.
  • updateTransaction: sets the transaction to its terminal state, then re-triggers processInstruction for the next hop. It fires either straight from processTransaction, or when the settlement webhook lands: bank-events drops it onto the same BankTransferServiceQ, and onEventReceived calls updateTransaction.

Internal, system-initiated transfers (like recurring investment debits) invoke createTransfer directly, with async: true where nobody needs to wait.

Every function pairs with its own healing job on an EventBridge rate() schedule, putting anything stuck in a non-terminal state back on the queue with exponential backoff. An end-of-day job checks against Green Dot for missed updates and status mismatches, and fixes what it finds.

If a hop still fails after 5 retries, its transaction goes VOID and the instruction FAILED. If an earlier hop has already moved money, nothing rolls back on its own: the instruction lands in SUPPORT_REQUIRED, and support can re-trigger the hop once Green Dot is back. That takes Green Dot being down for a day, or bad code. Logs carry specific keys, a Datadog monitor watches for them, and PagerDuty pages someone once a threshold is crossed, including for anything in SUPPORT_REQUIRED.

What it cost

A fully async, multi-hop system is genuinely harder to observe than a linear one: there’s no single call stack to trace. The trade-off was accepted deliberately, and covered by disciplined per-step logging so every hop’s state is traceable even though the execution path itself isn’t linear.

Where it landed

One codebase, one set of logic, handling both real-time and background transfers identically underneath. End-to-end transfer completes in under 5 seconds, excluding actual ACH settlement time. Immutable transaction records, one-directional state transitions, and active PagerDuty monitoring for anything that fails past the point self-healing can fix.