Latency vs Throughput
Concept
- Latency: The time it takes for a single message to travel from the sender to the receiver and back. (Measured in milliseconds,
ms). - Throughput: The total volume of data or number of requests a system can process within a specific time window. (Measured in Requests Per Second,
RPS, or Megabytes per second).
Mental Model: The Highway
Trade-Offs
- Often, optimizing for one degrades the other.
- Optimizing for Throughput: You might collect 1,000 log entries in memory, wait 5 seconds, and then write them all to the database in one massive batch. This achieves massive throughput (1,000 logs/sec), but terrible latency (the first log had to wait 5 seconds to be saved).
- Optimizing for Latency: You write every single log entry to the database the exact millisecond it arrives. Latency is incredibly low (1ms), but the database is overwhelmed by thousands of individual network connections, destroying your overall throughput capacity.
Real-World Usage
- Multiplayer Gaming (FPS): Strictly optimizes for Latency. If the latency (ping) is over 50ms, the game is unplayable. The actual amount of data being sent (throughput) is tiny (a few bytes containing X/Y coordinates).
- Netflix Video Streaming: Optimizes for Throughput. Once the video starts playing, a 500ms delay in fetching the next chunk of video doesn’t matter because the video player buffers it. But the system must be able to push gigabytes of data per second.
Interview Questions
Q: A user in Australia makes an API request to a database hosted in New York. The database query takes 1ms. Why does the user experience 250ms of latency?
A: Physical distance limits speed. The speed of light traveling through fiber-optic cables under the ocean takes roughly 100-150ms just to get from Australia to New York and back. You cannot fix this with faster code. To reduce latency, you must use a Content Delivery Network (CDN) or replicate the data to a data center geographically closer to the user (e.g., an AWS region in Sydney).
Q: We added 5 more app servers behind our load balancer. Did we improve Latency or Throughput?
A: We improved Throughput. The system can now handle 5x as many concurrent users. However, the time it takes for a single request to travel to the server, query the DB, and return (Latency) remains exactly the same.