Cross-border payments
Regulated money movement, without paying for two of everything.
A remittance platform growing across regions needed partner-grade compliance and predictable cost while traffic spiked around paydays and holidays.
- Company
- Aspora
- Stage
- Y Combinator W22
- Stack
- AWS, multi-region, Kubernetes, Datadog, Terraform
- Engagement
- Managed cost operations and compliance automation
- Agents
- Finly AI, Iris AI, Sage AI, ClearRisk
The situation
Aspora moves money across borders for a fast-growing diaspora customer base. Regulators and banking partners in each corridor expect controls, data handling rules and evidence, and the product has to stay up when everyone sends money on the same day.
To meet residency and latency needs, environments were duplicated per region. Traffic followed paydays and holidays, so capacity was set for the peaks. Datadog spanned every environment, and the on-demand AWS bill grew with each new corridor.
Growth made the pattern worse, not better. Each new corridor arrived with its own partner requirements, a new region or environment, and a set of monitors copied from the last one. Costs were tracked, but tracked spend and understood spend are different things, and nobody on a lean team owned the second.
What the audit found
5 things that were costing money or time.
Regional duplication with idle capacity
Each region carried full-size environments around the clock, including non-production. Peak sizing was right for a few days a month and wrong for the rest, and non-production environments in every region mirrored production down to the instance types.
AWSEverything on demand
A stable compute floor across regions ran on on-demand pricing with no commitments, the most expensive way to buy capacity that never turns off. The floor was easy to see once utilization was lined up across regions. Nobody had had the time to line it up.
AWSTransaction pipeline logs indexed in full
Payment pipelines emit a lot of structured logs. All of it was indexed at full retention across every region, most of it never searched. Investigations needed a few fields from a few services, and paid for all of them from all of them.
DatadogAlert noise across regions
Monitors copied per region paged for regional blips that customers never noticed, so real incidents competed with noise. The on-call engineer had learned which alerts to ignore, which is the most dangerous kind of tuning.
DatadogCompliance evidence by hand
Access reviews, key rotation and change approvals were done, but proving it to partners meant a manual collection exercise each time, repeated for each corridor's requirements.
ComplianceWhat changed
Agents did the work. Engineers approved it.
Commitments and schedules
Savings Plans matched to the steady-state floor, autoscaling bounds set from real peaks, and non-production environments scheduled per region. Payday and holiday peaks are handled by autoscaling headroom rather than permanent capacity.
Log routing and monitor consolidation
Exclusion filters and archive routing for pipeline logs with rehydration when an investigation needs them, and region-aware monitors that page on customer impact. The alerts the on-call engineer had learned to ignore were deleted on purpose, with a record of why.
Multi-region infrastructure as code
Regions expressed as one module with per-region variables, policy checks on every plan, and evidence for access, change and key rotation exported continuously. A new corridor is now a variables file and a pull request.
Vulnerability triage
CVEs correlated with what runs in each region, patch pull requests for the ones that reach production, and a record partners can review.
Results
What the next invoices showed.
The log and monitor changes went first and changed the on-call experience within days: fewer pages, each one meaning something, and investigations that rehydrate the logs they need rather than paying to keep everything hot. The Datadog bill followed the volume down.
On AWS, the commitments were the largest single change and the one that needed the most confidence. Thirty days of utilization lined up across regions showed a floor the platform had never dropped below, and the Savings Plans were sized to that floor, not to the peaks. Schedules took non-production environments offline outside working hours in each region. Peak days are now absorbed by autoscaling headroom that costs nothing when it is not in use.
For the partner conversations, the change was in kind rather than degree. Evidence for access reviews, key rotation and change approvals is exported as it is produced, so a request that used to start a collection exercise now ends with a link. Adding a corridor became an infrastructure change like any other: a variables file, a plan, an approval.
What we learned
Three things we'd tell a peer.
- Multi-region is a multiplier. Every default you never revisited is now paid for several times over.
- Commitments are safe once the floor is measured and reckless before. Measure first.
- Deleting an alert deliberately, with the reason recorded, is better tuning than an engineer learning to ignore it.
Timeline
How the engagement ran.
- Week 1
Read-only audit across regions
One role per account, Datadog read keys, thirty days of utilization and usage.
- Weeks 2 to 4
Logs, monitors, commitments
No-risk changes first, then commitments once the floor was confirmed, each behind an approval.
- Month 2 onward
Schedules, IaC and evidence
Regional schedules, the multi-region module, and continuous evidence export.
Same pattern in your bill?
Book 30 minutes with an engineer to scope a free, read-only audit.
More case studies: Coinshift, StartGlobal, Inc.
You'll pick a time on the next page.