Series: Building Backarch, engineering decisions from building backarch.com
Three cloud providers, 10 regions each, dozens of service categories, and pricing structures that have almost nothing in common. Most teams hardcode the numbers and pretend they're still accurate six months later.

Here's what it actually takes to keep them live, and the pricing shape distinction that trips everyone up the first time.
Three fetchers, three completely different APIs
The pricing cron lives in pricing_cron/run_pricing_cron.py, a standalone Python script with its own process, its own database connection, its own environment loading. Three fetchers run concurrently, each dispatched to a thread pool via loop.run_in_executor(), since none of them are natively async. Each speaks a completely different language.
AWS uses boto3's pricing.get_products(), a filter-based API requiring service codes like AmazonEC2, AmazonRDS, AWSQueueService. The response is a list of JSON blobs. The price lives three levels deep: priceDimensions is a dict keyed by an opaque ID, and pricePerUnit inside that is a dict keyed by currency. Parsing it correctly requires reading several pages of AWS documentation and then testing against the live API.
Azure has the cleanest API of the three, a public REST endpoint at prices.azure.com/api/retail/prices requiring zero authentication. But it returns paginated responses with a NextPageLink, and a full region fetch takes several seconds. You cannot call this per-request without destroying latency. Pre-population is the only viable option.
GCP uses the Cloud Billing Catalog API, authenticated via google-auth ADC. GCP SKU names are unstable, a query that returned results last quarter can return 0 rows after a catalog update, which is a real failure mode the write path has to account for (more on that below).
All three run through the same executor pattern, not a sync/async split, because boto3 has no async client and the simplest way to keep every fetcher's error handling uniform is to treat all of them as blocking calls:
# pricing_cron/run_pricing_cron.py
loop = asyncio.get_running_loop()
async def run(name: str, fn, catalog):
try:
rows = await loop.run_in_executor(None, fn, catalog)
return name, rows
except Exception as exc:
logger.exception("Fetcher | %s | FAILED", name)
return name, exc
results = await asyncio.gather(
run("aws", _fetch_aws, catalogs["aws"]),
run("azure", _fetch_azure, catalogs["azure"]),
run("gcp", _fetch_gcp, catalogs["gcp"]),
)Each run() call catches its own exceptions and returns them as data instead of letting them propagate, which is what lets one provider's failure sit alongside the other two's success in the same results list. If Azure's API is down at 02:00 UTC Sunday, the AWS and GCP rows still update.
AI model pricing is a separate mechanism entirely, not part of this weekly cron. pricing_cron/ai.py reads litellm.model_cost, a dict bundled with the litellm package covering 900+ models across a dozen providers, at import time, with no network call and no authentication. It's cheap enough to compute on demand rather than needing a scheduled refresh.
The pricing shape problem nobody documents
Every row, regardless of fetcher, lands in the same shape: (provider, service, region, instance_type, unit, price_usd). Getting there correctly requires understanding that cloud services use three fundamentally different pricing shapes.
Regional live SKU. The price genuinely varies by region. EC2 t3.micro costs $0.0104/hr in us-east-1 and $0.0124/hr in ap-southeast-1. Fetch the actual price per region. Ten target regions means ten API calls and ten rows per service type.
Global live SKU. The price is identical everywhere, it's an account-level charge, not a regional one. AWS Lambda request pricing, DynamoDB on-demand reads, Route 53 hosted zones, Cognito MAUs. If you naively fetch per-region, you get 10 identical rows. If you skip the stamp-across-regions step, the compare endpoint returns no result for 9 of your 10 target regions.
# pricing_cron/fetch_pricing_aws.py, global SKU pattern
lambda_request_price = _fetch_lambda_request_price()
rows = [
PricingRow(
provider="aws", service="lambda",
region=region, instance_type="requests",
unit="per-1M-requests", price_usd=lambda_request_price,
)
for region in ALL_TARGET_REGIONS # stamp the same price across all 10
]Geography-band SKU. CloudFront prices by traffic geography, US/Canada/Europe, APAC, South America, not by AWS deployment region. us-east-1 is a deployment region, not a traffic origin. The fetcher grabs the actual US/NA band rate and stamps it across us-east-1, us-west-2, and ca-central-1. Same pattern for GCP Cloud CDN and Azure CDN. It's an approximation, but it's the actual published geography-band rate.

Most pricing documentation only covers regional pricing, because that's the common case. The global and geography-band cases require reading the pricing FAQ pages, not the API docs.
A zero-row response is a failure, not a signal to improvise
Some services don't expose usable pricing data through their APIs on a given run. A catalog change, a rate-limited call, a transient GCP SKU rename, any of these can bring a fetcher back with zero rows instead of an exception.
The temptation is to treat that as a special case worth working around, maybe hardcode last quarter's known-good number so the row is never empty. The actual rule is the opposite: a zero-row result is logged and treated exactly like an exception, and the existing data for that provider is left untouched.
# pricing_cron/run_pricing_cron.py
if isinstance(result, Exception):
failed.append(provider)
logger.error("Fetcher | %s | FAILED, keeping old data in DB", provider)
elif not result:
failed.append(provider)
logger.error("Fetcher | %s | returned 0 rows, treating as failure, keeping old data", provider)
else:
succeeded[provider] = resultA stale-but-correct price from last week's successful run is a better default than a number a person typed in once and forgot to revisit. If every provider fails in the same run, nothing gets written at all, the cron exits non-zero and the existing table is untouched.
Replace-per-provider: the atomic refresh strategy
cloud_pricing_latest uses a replace-per-provider strategy: delete all rows for a provider, then batch-insert the fresh rows, committed together:
# pricing_cron/run_pricing_cron.py
async def _replace_provider(session: AsyncSession, provider: str, rows: list[dict]) -> int:
"""Delete existing rows for this provider, then batch-insert new ones."""
result = await session.execute(_DELETE_SQL, {"provider": provider})
logger.info("DB | %s | deleted %d existing rows", provider, result.rowcount)
inserted = 0
for i in range(0, len(rows), BATCH_SIZE):
batch = rows[i : i + BATCH_SIZE]
await session.execute(_INSERT_SQL, batch)
inserted += len(batch)
await session.commit()
return insertedEvery run produces a clean, consistent snapshot for that provider. This only runs for providers in succeeded, the ones that returned real rows, so a failed or empty fetch for one provider never touches that provider's existing rows, and never blocks the other two from updating.
There's also a staff-only endpoint for seeding a fresh environment without waiting for the next scheduled run:
POST /pricing/refresh
(multipart file upload, requires is_staff=True)
It doesn't trigger a live fetch. It reads an uploaded CSV, produced by the same cron's merge_to_csv.py output, and upserts it in batches, then clears the pricing read cache so the new numbers show up immediately instead of waiting out the TTL. You need it because cloud_pricing_latest is empty on a fresh deploy, and the next scheduled cron run might be days away.
Building the cost estimation feature at backarch.com meant every component node on the canvas needed a real price, hardcoded numbers from a spreadsheet six months old don't make a useful product feature. The part that took longest wasn't the fetchers, it was figuring out which of the three price shapes applied to each of the 30+ AWS services, then verifying it against a region where the published rate was known.
Live pricing data is not a nice-to-have for a cost estimation feature. It's the feature. The fetcher infrastructure to keep it live is the engineering cost of that claim.
