FAN YOURSS

系统设计[实战篇-Temperature Reading]

Step 1 - Understand the problem and establish design scope
Step 2 - Propose high-level design and get buy-in
Step 3 - Design deep dive
Step 4 - Wrap up

题目是大概有3B的daily active user,每个人都有一个手机,手机会读取天气信息,然后上传给服务器,服务器提供API可以query到每个城市当前的天气。

  1. Define Problem
  2. 定量分析,*trade off*选择,(这题focus的是减少数据存储量,cost optimization
  3. 设计数据结构
  4. 设计system diagram (Make sure it works)
  5. 分析reliability
  6. 找寻设计中的问题,以及可以优化的部分

Define Problem

需求在于三方面,1.收集来自用户的数据 2.aggregate数据 3.提供读API

需要clarify的点在于:

  1. 数据precision的要求?时间上,每个小时更新一次还是每分钟?位置上,City Level 还是Zip Code Level? (Fun fact, ~ 19k city in US, ~4M world wide)
  2. Aggregation function: average
  3. Read API需要maintain 所有的temperature history吗,还是只需要当前温度?

Problem & Scope

从3B DAU 的手机上每小时读取一次所在city的temp reading,并提供city level当前的气温数据,不需要历史记录。

3B DAU = 3000000000 user per day,

1 request per hour = 36000000000 request/day ~ 833333 rps

36000000000 * 256 byte = 9.216 TB per day

(数据量和rps提供了大概的感觉,这些cost不太值得记录所有的数据,同时因为只需要city level的数据,同一个city,重复的会很多)

High Level Design

Trade-Off 1

Device Ratelimiting vs Server Ratelimiting (Preferred)

Ratelimit on City (100 events per hour)

Ratelimit 完了以后,这个就变成了trivial problem.

用一个数据库记录所有的datapoint, hour, city, tmp

然后用grafana query就行了

这里完全可以简化设计,使用prometheus去log timeseries data,数据是ts + city + tempture

这里需要设计一下push pipeline,还是需要一个Events API Service去接收以及sample events (ratelimiting on Time_Hour + City < 100)

之后给prometheus scrape,scrape完以后就可以用prom query

Focus trade off analysis:

  • Store all events VS sampling (huge cost saving)
  • Client send all events VS service deriving data (prevent skewed data, but subject to geolocation failures)
  • Prometheus VS other (Simplicity, less customizability -> track history)

Failure

  • Geolocation failure (use client data as backup)
  • Rate limiter failure (fail open vs fail close, if lost data is not acceptable fail open)
  • event store failure (nothing can be done)