系统设计[实战篇-Temperature Reading]
Step 1 - Understand the problem and establish design scope
Step 2 - Propose high-level design and get buy-in
Step 3 - Design deep dive
Step 4 - Wrap up
题目是大概有3B的daily active user,每个人都有一个手机,手机会读取天气信息,然后上传给服务器,服务器提供API可以query到每个城市当前的天气。
- Define Problem
- 定量分析,*trade off*选择,(这题focus的是减少数据存储量,cost optimization
- 设计数据结构
- 设计system diagram (Make sure it works)
- 分析reliability
- 找寻设计中的问题,以及可以优化的部分
Define Problem
需求在于三方面,1.收集来自用户的数据 2.aggregate数据 3.提供读API
需要clarify的点在于:
- 数据precision的要求?时间上,每个小时更新一次还是每分钟?位置上,City Level 还是Zip Code Level? (Fun fact, ~ 19k city in US, ~4M world wide)
- Aggregation function: average
- Read API需要maintain 所有的temperature history吗,还是只需要当前温度?
Problem & Scope
从3B DAU 的手机上每小时读取一次所在city的temp reading,并提供city level当前的气温数据,不需要历史记录。
3B DAU = 3000000000 user per day,
1 request per hour = 36000000000 request/day ~ 833333 rps
36000000000 * 256 byte = 9.216 TB per day
(数据量和rps提供了大概的感觉,这些cost不太值得记录所有的数据,同时因为只需要city level的数据,同一个city,重复的会很多)
High Level Design
Trade-Off 1
Device Ratelimiting vs Server Ratelimiting (Preferred)
Ratelimit on City (100 events per hour)
Ratelimit 完了以后,这个就变成了trivial problem.
用一个数据库记录所有的datapoint, hour, city, tmp
然后用grafana query就行了

这里完全可以简化设计,使用prometheus去log timeseries data,数据是ts + city + tempture
这里需要设计一下push pipeline,还是需要一个Events API Service去接收以及sample events (ratelimiting on Time_Hour + City < 100)
之后给prometheus scrape,scrape完以后就可以用prom query

Focus trade off analysis:
- Store all events VS sampling (huge cost saving)
- Client send all events VS service deriving data (prevent skewed data, but subject to geolocation failures)
- Prometheus VS other (Simplicity, less customizability -> track history)
Failure
- Geolocation failure (use client data as backup)
- Rate limiter failure (fail open vs fail close, if lost data is not acceptable fail open)
- event store failure (nothing can be done)