The Hidden Cost of Cloud Market Data: The Compute Spend Financial Institutions Aren’t Tracking

By Alistair Brooker, VP General Manager, Calero

Financial institutions have become increasingly adept at managing the direct costs of market data. Market Data and Procurement teams track vendor subscription fees line by line, and IT groups understand the per-user licensing costs tied to platforms like Snowflake or Databricks. But a growing expense often slips through the cracks: the cloud compute consumed as users query, stream, download and analyze the massive volumes of data those subscriptions unlock. As market data workflows shift onto cloud-native architectures, compute charges from hyperscalers such as AWS, Google Cloud and Microsoft Azure can quietly outpace expectations, and unlike subscription fees, they are notoriously difficult to attribute to any one user, desk or business unit.

How cloud created a blind spot

Market Data most commonly arrives through dedicated feeds and on-premises infrastructure, where cost is largely fixed and predictable: a feed handler, a set number of terminals, a known hardware or web-based footprint. Cloud delivery changed that equation. Data now lands in a warehouse or lakehouse, where quants, risk teams and trading desks run ad hoc queries, backtests and large-scale historical analyses against it. Every one of those actions consumes compute, and compute is billed by usage rather than by seat.

Alistair Brooker

Cloud delivery brought real benefits: elastic scaling, faster time to insight, and the ability to run analysis that would have been impractical on-premises. But it also decoupled cost from the thing organizations were used to measuring. A subscription invoice tells finance exactly what was paid for data access. A cloud bill tells them only that a large number of compute credits were consumed somewhere across the organization, by workloads that are rarely tagged, labelled or attributed with enough granularity to trace back to a source.

Why allocation is so hard

The data is there. The problem is making sense of it across the organization. Query logs, warehouse credit consumption and job-level billing all exist inside the platform, but they are rarely organized in a way that maps cleanly to organizational structure. A single dataset might be queried by multiple teams for different purposes, a scheduled job might run under a shared service account, and a research analyst’s exploratory query can consume as much compute as a production pipeline running all day. Without consistent tagging discipline and a system to interpret it, IT and finance teams are left staring at an aggregate number with no reliable way to break it down by user, desk or cost center.

That makes even basic questions about spend surprisingly difficult to answer. Who is actually driving the cost? Which workloads are responsible? And where is there room to optimize? Teams can’t rationalize inefficient queries they don’t know exist, and they can’t renegotiate data entitlements or compute tiers based on actual usage patterns they can’t see.

Governance is catching up, slowly

FinOps practices, which matured around general cloud infrastructure spend, are now being pulled into this more specialized corner of the business. The basic principles still apply: workloads need to be tagged at the source, shared costs need to be allocated consistently, and business units need enough visibility into their own consumption to have a reason to manage it. What’s different with market data is the added layer of vendor entitlement and licensing complexity sitting on top of the infrastructure cost, which means allocation models built purely for generic cloud spend often miss the mark. That makes the question bigger than simply who used the most compute. Effective governance has to connect the data an organization has licensed with the teams actually using it, while also showing what that usage is costing in the cloud. Those pieces have traditionally lived in different systems, making it difficult to see the full picture in one place.

Extending that visibility into cloud compute usage means organizations can analyze how market data is actually consumed across Snowflake, Databricks and the underlying hyperscaler environments, and attribute that consumption back to the teams and business units generating it. Rather than treating vendor fees and compute spend as two disconnected line items, Calero brings them into a single view, giving finance and IT the accountability and cost control that cloud-native market data delivery has so far made difficult to achieve.

spot_img
spot_img

Subscribe to our Newsletter