How does BigQuery separate storage and compute?
BigQuery is a serverless, columnar data warehouse built on a separation of storage and compute.
- Storage: data lives in Capacitor, a columnar format backed by Colossus distributed storage. You pay for bytes stored, with long-term automatic discounts after 90 days.
- Compute: queries run on a shared pool of Dremel slots. You can use on-demand pricing, billed per byte scanned, or flat-rate and slot reservations for predictable workloads.
- This separation means storage scales independently of query capacity, and you can query data without provisioning clusters.
SELECT user_id, COUNT(*) AS events
FROM `proj.analytics.events`
WHERE event_date = CURRENT_DATE()
GROUP BY user_id
ORDER BY events DESC
LIMIT 100;
Optimise cost by partitioning on date, clustering on common filters, and selecting only needed columns. Avoid SELECT *, and check the query plan before running at scale.