Scalability

⭐ Interview Importance: HIGH
⏱️ Revision Time: 4 min

Concept

Scalability is the ability of a system, network, or process to handle a growing amount of work, or its potential to be enlarged to accommodate that growth. In System Design, we generally scale in two ways: Vertical Scaling (scaling “up”) and Horizontal Scaling (scaling “out”).

Architecture Diagram

Trade-Offs

Vertical Scaling (Scaling Up)

  • Definition: Adding more resources (CPU, RAM) to a single existing node.
  • Pros:
    • Extremely simple. No code changes required.
    • Less administrative overhead (managing 1 server vs 100).
    • Inter-process communication is fast (everything is in memory).
  • Cons:
    • Hardware Limit: There is a hard physical limit to how big a single server can get.
    • Single Point of Failure (SPOF): If that one massive server crashes, your entire application goes offline.
    • Costly. Enterprise-grade monolithic servers are exponentially more expensive than commodity hardware.

Horizontal Scaling (Scaling Out)

  • Definition: Adding more nodes/machines to the system.
  • Pros:
    • Infinite scalability. Just keep adding cheap commodity servers.
    • Built-in redundancy. If one server dies, the load balancer reroutes traffic to the others.
  • Cons:
    • High complexity. Requires Load Balancers, Service Discovery, and network configuration.
    • State Management: You can no longer store user sessions in memory (since the next request might hit a different server). You are forced to use a distributed cache like Redis.
    • Network latency between nodes.

Real-World Usage

  • Startups: Start with Vertical Scaling. Run the DB and App on a single DigitalOcean Droplet. When you max it out, buy the $80/mo tier. Do not over-engineer a microservice cluster for 100 users.
  • Databases: Relational Databases (MySQL/PostgreSQL) are traditionally much easier to scale vertically than horizontally due to ACID constraints. NoSQL databases (Cassandra/DynamoDB) are designed explicitly for horizontal scaling via sharding.

Interview Questions

Q: You have a single server running a Node.js application. CPU utilization hits 100%. What is your first step to scale?
A: Because Node.js is single-threaded, buying a server with 16 CPUs (Vertical Scaling) won’t help if you run it normally; Node will still only use 1 CPU. You must first use the Node.js cluster module or a process manager like PM2 to spawn 16 worker processes on that single machine. If that maxes out, then I would introduce a Load Balancer and scale horizontally across multiple machines.

Q: How does horizontal scaling affect user sessions?
A: If Server 1 authenticates the user and stores their session in its local RAM, a subsequent request routed to Server 2 will fail because Server 2 doesn’t know who the user is. This is solved by either using Sticky Sessions (the Load Balancer always routes the same user to the same server) or by making the servers stateless and moving session data to an external, centralized datastore like Redis.