How Many Users Can a 1-Core, 2GB RAM Server Actually Handle? We Stress-Tested It to the Breaking Point

The CyberSec Guru

How Many Users Can a 1 vCPU 2GB RAM Server Handle

If you like this post, then please share it:

Buy me A Coffee!

Support The CyberSec Guru’s Mission

🔐 Fuel the cybersecurity crusade by buying me a coffee! Why your support matters: Zero paywalls: Keep the main content 100% free for learners worldwide.

“Your coffee keeps the servers running and the knowledge flowing in our fight against cybercrime.”☕ Support My Work

Buy Me a Coffee Button

A hands-on load testing experiment reveals that a $6/month VPS running Nginx, Node.js, and PostgreSQL simultaneously can sustain thousands of concurrent users — but only until the CPU saturates. Here’s the complete technical breakdown, the exact failure thresholds, and the single caching optimization that doubled capacity without touching hardware.

The Premise: One Server, Everything on It, Zero Scaling

There’s a persistent myth in software engineering circles — perpetuated by cloud vendor marketing and over-engineered architecture diagrams on social media — that you need Kubernetes clusters, auto-scaling groups, managed databases, CDN layers, and at least three microservices before you can serve a single production user. The reality, as this experiment demonstrates with brutal empirical clarity, is far more forgiving for the vast majority of applications on the internet.

The experiment in question is deceptively simple in its setup but remarkably rigorous in its execution. A single DigitalOcean Basic Droplet — one shared vCPU, 2 GB RAM, 50 GB SSD storage, running Ubuntu — hosts the entire application stack. No load balancer in front. No Redis cluster behind. No externally managed PostgreSQL instance sitting in a separate VPC. No Docker orchestration layer. No CDN distributing static assets. The web server (Nginx), the application runtime (Node.js), and the relational database (PostgreSQL) all compete for the same single CPU core and the same 2 gigabytes of memory.

The question being answered isn’t theoretical. It isn’t derived from CPU cycle calculations or memory allocation spreadsheets. It’s empirical: spin up a realistic application, model realistic user behavior, and hammer the server with thousands of concurrent virtual users until the latency thresholds are violated and the system effectively collapses. Then, apply targeted optimizations and measure how much additional capacity is unlocked.

This is the kind of benchmarking that most engineering teams skip entirely until production traffic forces the conversation — usually at 2 AM during an incident. Having concrete numbers before you need them is not just good engineering practice; it’s operational insurance.

Test Environment and Infrastructure Configuration

The Application Server

The server under test is DigitalOcean’s cheapest Basic Droplet tier. For those unfamiliar with cloud VPS pricing structures, this is approximately $6/month at the time of testing. The specifications are:

  • CPU: 1 shared vCPU (not dedicated — this distinction matters and we’ll return to it)
  • RAM: 2 GB
  • Storage: 50 GB SSD
  • Operating System: Ubuntu (default DigitalOcean image)
  • Location: New York data center

The “shared vCPU” designation is critical to understand. On DigitalOcean’s Basic tier, your virtual machine is allocated a slice of a physical CPU core that is time-shared with other customers’ VMs on the same physical host. This introduces the concept of CPU steal time — the percentage of time your VM wanted CPU cycles but didn’t get them because the hypervisor was servicing another tenant’s workload. AWS calls this “CPU credit” depletion on burstable instances (t2/t3 families); DigitalOcean exposes it more directly.

In this experiment, CPU steal time was monitored throughout all test runs and consistently measured between 0% and 2%. Statistically insignificant. This aligns with what most users report on DigitalOcean’s Basic tier during non-peak hours: the oversubscription ratio on modern hypervisors (KVM-based, in DigitalOcean’s case) rarely produces measurable contention for sustained single-core workloads unless you’re on an extraordinarily unlucky physical host.

The Load Generation Server

Load Testing Topology
Load Testing Topology

Generating realistic traffic load requires its own dedicated compute resources. Running a load generator on the same machine being tested would contaminate results — the load generator’s own CPU and memory consumption would steal resources from the application under test, producing artificially low throughput numbers and making it impossible to determine whether failures were caused by the application hitting its limits or by the test harness starving it.

A second DigitalOcean Droplet was provisioned specifically for traffic generation:

📬 Stay Ahead of Cyber Threats

Get the latest cybersecurity news, critical vulnerabilities, threat intelligence, tutorials, and exclusive giveaways delivered straight to your inbox. No spam. Unsubscribe anytime.

Subscribe to the Newsletter →
  • CPU: 4 dedicated vCPUs
  • RAM: 8 GB
  • Location: Same New York data center (critical for minimizing network latency between generator and target)
  • Rental duration: Approximately 4 hours total (cost: under $1)

Co-locating both servers in the same data center reduces inter-server latency to sub-millisecond levels. This is intentional. The goal is to measure the application server’s intrinsic performance ceiling, not to measure the compound effect of application processing time plus 40 milliseconds of transatlantic network round-trip time. If you’re benchmarking your own infrastructure, replicate this: put your load generator in the same availability zone or data center as your target.

The Software Stack

The architecture is deliberately minimal:

Single-Server Architecture
Single-Server Architecture

All three components run on the same host. Nginx accepts incoming HTTP connections, parses headers, performs basic request routing, and proxies the request to the Node.js process (typically listening on a local port like 3000). Node.js handles application logic — authentication verification, request validation, business rules — and issues SQL queries to PostgreSQL. PostgreSQL executes the query against the dataset and returns results up the chain.

No connection pooling middleware like PgBouncer sits between Node and Postgres. No process manager like PM2 is spawning multiple Node workers (single process, single thread event loop). No Redis or Memcached layer intercepts repeated queries. This is the “simplest thing that could possibly work” architecture, and that’s precisely the point.

The Application: A Realistic Social Media Backend

Why Not “Hello World”?

Benchmarking against a GET /health endpoint that returns {"status": "ok"} tells you nothing about real-world capacity. It exercises Nginx’s static response path, maybe touches Node’s event loop once, and never opens a database connection. It’s the computational equivalent of measuring how fast a car goes in neutral.

The application built for this test is a small social media backend — functionally similar to a stripped-down Twitter or Instagram API. It supports four operations:

EndpointMethodOperation
/feedGETRetrieve a paginated feed of posts
/post/:idGETRetrieve a single post by ID
/post/:id/likePOSTLike a specific post
/postPOSTCreate a new post

This is intentionally representative of read-heavy web applications. The ratio of reads to writes (feed loads and post views versus likes and new posts) mirrors real social platforms where read operations outnumber writes by roughly 10:1 or more.

The Dataset

Testing against an empty database would be meaningless — PostgreSQL would return results from memory-mapped buffer pages almost instantly, giving no indication of how the system behaves under realistic data volumes. The database was pre-seeded with:

  • 50,000 user accounts
  • 500,000 posts
  • 2,000,000+ likes (post-user relationships)

Total database size on disk: approximately 360 MB. This is small by production standards (a mid-sized social platform’s Postgres instance might be 500 GB to several terabytes), but it’s large enough that PostgreSQL must perform actual index lookups, join operations, and buffer management rather than serving everything from the OS page cache trivially.

The database was configured with reasonable parameters for a 2 GB RAM environment — shared_buffers set to approximately 25% of available RAM (512 MB), work_mem tuned for the query complexity, and effective_cache_size reflecting the total memory available to the OS for file caching. These aren’t exotic tuning parameters; they’re what any competent engineer would configure on a small VPS.

Similarly, indexes were added where any reasonable engineer would add them: primary keys on all tables, foreign key indexes on posts.user_id, likes.post_id, and likes.user_id, and a composite index on the feed query’s sort columns. This isn’t a sabotaged benchmark. The application is configured the way a junior-to-mid-level engineer would ship it.

Load Testing Methodology: Modeling Real Human Behavior

Why Raw Requests Per Second Is a Misleading Metric

You could point a tool like ab (Apache Bench) or wrk at the server and fire 10,000 simultaneous requests at a single endpoint. If they all return HTTP 200, you could technically claim the server “handled” 10,000 users. But that’s not how humans interact with software. A human loads a page, reads it, thinks, scrolls, clicks something, reads that, maybe types a comment. There are pauses. There’s randomness. There’s session state.

The load testing framework used here is Grafana k6 — an open-source load testing tool written in Go that uses JavaScript for test scripting. k6 models load in terms of Virtual Users (VUs), where each VU executes a scripted behavioral loop with configurable think-time pauses between actions.

The Virtual User Behavioral Loop

Virtual User Load Testing Loop
Virtual User Load Testing Loop

Each virtual user in this experiment executes the following sequence independently:

  1. Load the feed → 1 GET request to /feed
  2. Think time: Random pause between 3 and 7 seconds
  3. Open a post → 1 GET request to /post/:id (the post is selected from the feed results)
  4. Think time: Random pause between 3 and 8 seconds
  5. Conditional action: 15% probability of liking the post (POST to /post/:id/like), 2% probability of creating a new post (POST to /post)
  6. Think time: Random pause between 5 and 15 seconds
  7. Loop back to step 1

Every virtual user runs this loop independently with randomized timing. With 2,500 concurrent VUs, you do not have 2,500 requests arriving at the server in the same millisecond. You have 2,500 independent agents at various points in their loop, generating a statistically distributed stream of requests.

The Math on Per-User Load

The average loop duration is approximately 20 seconds. Each loop generates slightly over 2 requests (one feed load + one post view + the probabilistic like/create). This means each active user generates roughly 0.1 requests per second.

This is a critical number to internalize. It means that 2,500 concurrent active users produce approximately 250 requests per second of aggregate load. That’s not a lot in raw throughput terms. But the server must handle each request with correct authentication, database queries, response serialization, and network I/O — all on one CPU core.

Authentication

Every virtual user was assigned a unique pre-seeded user account with a pre-signed JWT (JSON Web Token) for authentication. This prevents the unrealistic scenario of all requests appearing to come from the same user, which would skew caching behavior, database query patterns, and any per-user rate limiting logic. Each VU is effectively a distinct authenticated person.

Defining “Success”: The Latency and Error Thresholds

Before any load was applied, three hard thresholds were defined. If any of these were violated during a test run, that user count was considered a failure:

MetricThresholdMeaning
P95 Latency< 500 ms95% of all requests must complete in under half a second
P99 Latency< 1,000 ms99% of all requests must complete in under one second
Error Rate< 1%Fewer than 1 in 100 requests can return a non-2xx status

These are not arbitrary. A P95 of 500 milliseconds aligns with widely cited web performance research (Google’s RAIL model, Nielsen Norman Group’s usability studies) indicating that users perceive responses under 500 ms as “instant” and begin noticing degradation above 1 second. A P99 under 1 second ensures that even the slowest tail requests — typically caused by garbage collection pauses, database lock contention, or TCP retransmissions — don’t exceed the threshold where users start abandoning sessions.

The sub-1% error threshold accounts for the reality that at high load, some connections will time out, some TCP handshakes will fail, and some requests will hit transient resource exhaustion. But if more than 1% of your users are getting errors, your service is effectively down for a meaningful fraction of its audience.

It’s worth emphasizing: the server could technically “handle” 10,000 users if it simply queued every request and responded after 45 seconds. The thresholds prevent this gaming. They enforce that the server must respond quickly enough to be usable, not just eventually.

The Load Ramp: From 10 Users to System Failure

Server Latency Under Increasing Load
Server Latency Under Increasing Load

Phase 1: Sanity Check (10–200 Users)

The test began conservatively. At 10 concurrent users, the system responded perfectly. At 50, 100, and 200 users, no degradation was observable. Latency remained in the low single-digit milliseconds. CPU utilization was negligible. This phase serves as a validation that the test harness, network configuration, and application deployment are all functioning correctly before applying serious load.

Phase 2: 1,000 Concurrent Users — The Server Doesn’t Care

At 1,000 simultaneous active users, the server was handling approximately 83 requests per second. The latency profile:

  • Median (P50): ~5 ms
  • P95: 19 ms
  • P99: 50 ms

These numbers are extraordinary for a single-core machine running an application server and a database simultaneously. A P99 of 50 milliseconds means that even the slowest 1% of requests — the ones that hit a cold database page, or caught Node’s event loop mid-callback, or waited on a TCP write buffer — completed in one-twentieth of a second. The server was essentially idle relative to its capacity.

Phase 3: 2,000 Concurrent Users — Tail Latency Emerges

Doubling to 2,000 users pushed throughput to 166 requests per second. The median barely moved:

  • Median (P50): ~6 ms (up from 5 ms — imperceptible)
  • P95: 161 ms (up from 19 ms — 8.5x increase)
  • P99: 276 ms (up from 50 ms — 5.5x increase)

This is the first visible stress fracture. The median user experience is unchanged. The average user loading the feed or opening a post notices nothing different. But the tail — the unlucky 5% or 1% of requests that hit the system at a moment of peak contention — are getting dramatically slower. This is the classic signature of a system approaching CPU saturation: the queue starts forming, and the requests at the back of the queue wait disproportionately longer.

In queueing theory terms, this is the transition from low utilization (where latency ≈ service time) to moderate utilization (where latency ≈ service time + queuing delay). The M/M/1 queue model predicts that as utilization (ρ) approaches 1, mean waiting time grows as ρ/(1-ρ), which is non-linear and accelerates sharply near saturation.

Phase 4: 4,000 Concurrent Users — Total Collapse

At 4,000 users, the system was catastrophically overloaded. Throughput plateaued at approximately 262 requests per second — barely higher than the 2,000-user case, indicating the CPU was fully saturated and couldn’t process additional requests faster regardless of how many were queued.

  • Median (P50): 2.7 seconds
  • P95: 4.5 seconds
  • P99: ~13 seconds

The median user is now waiting nearly three seconds for a response. This is unusable. Users will abandon the page, refresh, generate duplicate requests, and make the situation worse. The P99 at 13 seconds means that 1% of requests are taking longer than most users will wait before closing the tab entirely.

Phase 5: Binary Search for the Exact Threshold

Rather than testing 3,000, 3,500, 3,750, etc. sequentially, the experiment used a binary search approach:

  • 3,000 users: Fails (P99 exceeds threshold)
  • 2,500 users: Passes
  • 2,750 users: P95 passes, but P99 exceeds 1 second → Fails

The confirmed maximum sustainable load: 2,500 concurrent active users.

Phase 6: Confirmation Runs

Because the test involves randomized think times and probabilistic actions, there’s inherent variance between runs. To confirm that 2,500 wasn’t a lucky result, three independent test runs were executed at that level. All three passed with consistent results:

  • Median latency: ~11 ms
  • P95 latency: ~288 ms
  • P99 latency: ~775 ms
  • Errors: Zero across all three runs
  • Total requests served per run: ~100,000

Three runs, 300,000 total requests served, zero errors, all within latency thresholds. This is not a fluke. The server can sustain 2,500 concurrent active users under this workload pattern.

What Actually Broke: The CPU Bottleneck

The failure mode was not memory exhaustion. It was not disk I/O saturation. It was not PostgreSQL connection limits or lock contention. It was not network bandwidth.

It was the CPU.

At 2,500 users (the passing threshold), CPU utilization was approximately 80%. At 3,000 users, it climbed into the 90% range. At 4,000 users, it was pinned at 99%.

This makes architectural sense. On a single-core machine running Nginx, Node.js, and PostgreSQL simultaneously, every request requires:

  1. Nginx worker process: TCP accept, HTTP parse, header inspection, proxy_pass setup
  2. Node.js event loop: JWT verification (cryptographic signature check), request routing, business logic, SQL query construction, response serialization
  3. PostgreSQL backend process: Query parsing, planning, index scan/heap fetch, row assembly, result return
  4. Context switches between all three processes competing for the same core

The JWT verification alone is computationally non-trivial. Verifying an HMAC-SHA256 or RSA signature for every request consumes CPU cycles. At 250 requests per second, that’s 250 cryptographic verifications per second on a single shared vCPU.

Once CPU hits 100%, requests cannot be processed faster. They queue in the kernel’s TCP accept queue, in Nginx’s upstream connection pool, in Node’s event loop callback queue, and in PostgreSQL’s process scheduling. Each layer of queueing adds latency multiplicatively. The system doesn’t crash — it just gets slower and slower until timeouts trigger and clients give up.

CPU Saturation and Queueing Latency
CPU Saturation and Queueing Latency

The CPU Steal Time Caveat

Because this is a shared vCPU tier, another customer’s VM on the same physical host could theoretically consume CPU cycles that would otherwise be available to this workload. The experiment measured steal time throughout all runs: 0–2%, statistically insignificant. If you’re running this test on your own infrastructure and see higher steal time (>5%), your results will be degraded by noisy-neighbor effects, and you should either re-test during off-peak hours or upgrade to a dedicated CPU instance.

Raw Throughput Validation: Hitting the Feed Endpoint Directly

To isolate maximum raw throughput from the behavioral modeling overhead, the experiment also tested the feed endpoint in isolation — no think times, no probabilistic actions, just continuous GET requests as fast as the load generator could produce them.

Result: approximately 280 requests per second before latency thresholds were violated.

This aligns precisely with the realistic user test. At 2,500 concurrent users generating ~0.1 requests/second each, the aggregate load was approximately 232 requests per second — roughly 83% of the server’s maximum raw throughput. The remaining 17% headroom explains why 2,500 passed but adding a few hundred more users pushed the system past its breaking point.

This cross-validation is important. It confirms that the behavioral model isn’t artificially inflating or deflating the results. The numbers are internally consistent.

Optimization Round 1: In-Process Caching in Node.js

The Observation

With the baseline established at 2,500 users, the next question was: without changing hardware, how much more capacity can be extracted?

The most obvious source of repeated work was the feed endpoint. In this simplified application, every user loads the same feed (in a real social media product, feeds are personalized, but they’re still heavily cached at various layers). Thousands of virtual users were executing the same SQL query, fetching the same rows from PostgreSQL, serializing the same JSON response — over and over, hundreds of times per second.

The Fix

A simple in-memory cache was added inside the Node.js process. The first request to /feed executes the database query and stores the result in a JavaScript Map or object with a short TTL (time-to-live). Subsequent requests within that TTL window return the cached response without ever touching PostgreSQL.

No Redis. No Memcached. No external cache service. Just a variable in Node’s heap memory. On a single-server architecture, this is perfectly valid — there’s no cache invalidation coordination problem because there’s only one application process.

The Result

The realistic user capacity jumped from 2,500 to 4,000 concurrent users. Same hardware. Same database. Same Nginx configuration. One Map.set() call in the application code.

To illustrate the magnitude of this change, here’s the 4,000-user test with and without caching:

MetricWithout CacheWith Node.js Cache
P95 Latency~4,500 ms~260 ms
P99 Latency~13,000 ms~1,000 ms
StatusCatastrophic failurePasses thresholds

Same 4,000 users. Same hardware. One order-of-magnitude improvement in tail latency from a single caching layer.

The Trade-Off

Caching introduces staleness. A cached feed response might be 5, 10, or 30 seconds old by the time the 500th user receives it. For a social media feed, this is usually acceptable — Twitter’s timeline doesn’t need to reflect a new post within 2 seconds of its creation for every reader. But caching makes debugging harder. When a user reports “I posted something but it’s not showing up,” the answer might be “it’s in the cache, wait 15 seconds.” This trade-off between performance and data freshness is a fundamental systems design decision.

Optimization Round 2: Moving the Cache to Nginx

The Observation

After the Node.js cache was implemented, every request still traversed the full path: Nginx → Node.js → (check cache) → return. Even for cached responses, Nginx was opening a connection to Node, Node was checking its in-memory Map, serializing the response, and sending it back up. The Node.js event loop was still being engaged for every single request, even when no actual computation was needed.

The Fix

Nginx has built-in proxy caching capabilities. By configuring proxy_cache_path, proxy_cache, and appropriate Cache-Control headers, Nginx itself can store the response from the upstream (Node.js) and serve subsequent identical requests directly from its own cache — without ever forwarding the request to Node.js or touching the application process.

This moves the cache one layer closer to the client. The principle in systems design is: serve data from as close to the user as possible. The hierarchy from farthest to closest:

  1. Database (PostgreSQL) — farthest, slowest
  2. Application server memory (Node.js heap)
  3. Reverse proxy cache (Nginx proxy_cache)
  4. CDN edge nodes (Cloudflare, CloudFront, Fastly)
  5. Client-side cache (browser, service worker, mobile app local storage) — closest, fastest

Moving from layer 2 to layer 3 eliminated all Node.js processing for cached feed requests. No event loop engagement. No JavaScript execution. No JSON serialization. Nginx reads the cached response from its own memory-mapped file or shared memory zone and writes it directly to the client socket.

The Result

The raw throughput of the feed endpoint went from:

  • Original (no cache): ~280 req/s
  • Node.js cache: ~1,200 req/s
  • Nginx cache: >10,000 req/s
Different Caching Strategies
Different Caching Strategies

The Nginx-cached throughput exceeded 10,000 requests per second, and the experiment’s load generator became the bottleneck before the application server did. The 4 vCPU / 8 GB load generation machine simply couldn’t produce HTTP requests faster than that.

With the realistic virtual user model (think times, probabilistic actions, mixed endpoints), the capacity increased from 4,000 to approximately 5,200 concurrent active users before the P95/P99 thresholds were violated.

Final passing metrics at ~5,200 users:

  • P95 latency: ~160 ms
  • P99 latency: ~474 ms
  • Error rate: 0%

The Final Tally: What One Server Can Do

ConfigurationMax Concurrent Active UsersApprox. Requests/Second
Baseline (no caching)2,500~232
+ Node.js in-memory cache4,000~400
+ Nginx proxy cache~5,200~520

All on 1 shared vCPU, 2 GB RAM, $6/month, running Nginx + Node.js + PostgreSQL on the same host.

And remember: these are concurrent active users — people actively making requests at the same instant. Depending on your application’s traffic pattern, 2,500 concurrent users could translate to:

  • 10,000–25,000 daily active users if your traffic is spread over 8–12 hours with a peak-to-average ratio of 3–4x
  • 50,000+ registered users if only 5–10% are active at any given time
  • More if your traffic is geographically distributed across time zones, flattening the peak
Final Capacity Comparison
Final Capacity Comparison

The “$6/month server can’t handle real traffic” narrative is, for the majority of small-to-medium web applications, simply incorrect.

Technical Analysis: Why CPU Was the Bottleneck (And Not RAM, Disk, or Network)

Why Not RAM?

2 GB is tight, but the working set for this workload fit within it. PostgreSQL’s shared_buffers at 512 MB, the OS page cache holding the 360 MB database file, Node.js heap at perhaps 100–200 MB, Nginx’s memory footprint at 20–50 MB, and the kernel’s own overhead — all of this fits within 2 GB without triggering swap. If the system had been swapping, latency would have shown a very different signature: sudden spikes of 100–500 ms corresponding to page faults and disk reads for swapped pages. That pattern was not observed.

Why Not Disk I/O?

The 50 GB SSD provides more than sufficient IOPS for this workload. PostgreSQL at 232 requests/second, with indexed queries against a 360 MB dataset, is not generating significant random I/O. The data fits almost entirely in the OS page cache after the first few queries warm it up. Write operations (likes, new posts) are minimal relative to reads and are buffered by PostgreSQL’s WAL (Write-Ahead Log) before being flushed to disk.

Why Not Network?

The responses in this workload are small JSON payloads — a feed of 20 posts might be 5–10 KB. At 280 requests/second, that’s roughly 2–3 MB/s of outbound traffic. Even a 100 Mbps network interface (the minimum on most VPS providers) can handle 12.5 MB/s. Network was never close to saturation.

Why CPU Specifically?

Every request, even a “simple” cached feed load, requires:

  • TCP/IP stack processing in the kernel (interrupt handling, socket buffer management)
  • Nginx worker process: epoll event notification, HTTP request parsing, header matching, upstream connection management
  • Context switches between kernel space and user space, between Nginx and Node processes
  • Node.js event loop: callback scheduling, V8 JIT-compiled JavaScript execution, JWT cryptographic verification (HMAC-SHA256 involves hashing operations), JSON serialization
  • PostgreSQL backend: query parser, planner/optimizer, executor, buffer pool lookup, response formatting
  • Inter-process communication: Nginx → Node via TCP on localhost, Node → PostgreSQL via Unix socket or TCP

All of this is CPU work. On a single core, it’s serialized. There’s no parallelism. The event loop is single-threaded. Nginx typically runs one worker per core (so one worker here). PostgreSQL forks one backend process per connection, but they all compete for the same core.

At ~250 requests/second, with perhaps 4–8 milliseconds of CPU time per request across all processes, you’re consuming 1,000–2,000 ms of CPU time per second — which is the entire capacity of one core. The math checks out.

Practical Implications for Engineers and Founders

When This Architecture Is Sufficient

If you’re building:

  • A SaaS product in its first 6–18 months
  • An internal tool for a team of 50–200 people
  • A content site with moderate traffic
  • An API serving a mobile app with a few thousand DAU
  • A side project or MVP

…a single $6–$12/month VPS running Nginx + your runtime + PostgreSQL is not just “enough.” It’s likely over-provisioned for your actual traffic. The premature optimization into microservices, managed Kubernetes, and multi-region deployments is, for most startups, a waste of engineering time that should be spent on product development.

When You Need More

This architecture breaks down when:

  • Your per-request CPU cost is high (image processing, video transcoding, AI inference, complex report generation)
  • Your database grows beyond what fits in RAM, causing disk I/O to dominate latency
  • Your write volume is high enough that PostgreSQL WAL writes become a bottleneck
  • You need zero-downtime deployments (a single server means you must stop the service to deploy)
  • Your traffic has extreme spikes (10x normal load during events, product launches, viral moments)

How to Test Your Own System

The methodology here is reproducible and straightforward:

  1. Profile your actual user behavior. Look at your analytics or access logs. What endpoints do users hit? In what order? With what think times between actions? What’s the read/write ratio?
  2. Script the behavior in k6. Grafana k6 is free, open-source, and well-documented. Most AI coding assistants (Claude, Copilot, Cursor) can generate a k6 script from a plain-English description of your user flow.
  3. Define your failure criteria. P95 latency, P99 latency, error rate. Write them down before you start the test. Don’t move the goalposts after.
  4. Generate load from a separate machine. Never load-test from the same server you’re testing. Use a separate VPS in the same data center. Rent it for a few hours; it’ll cost less than a coffee.
  5. Ramp up gradually. Start at 10 users. Double. Find the failure point. Binary search back to find the exact threshold.
  6. Monitor system resources during the test. htop, vmstat, iostat, sar, PostgreSQL’s pg_stat_activity. Know what is breaking, not just that something is breaking.
  7. Run confirmation tests. Once you find your threshold, run it three times. Variance is real. Make sure your result is reproducible.
Caching Hierarchy
Caching Hierarchy

The Bigger Picture: What This Experiment Actually Demonstrates

Beyond the specific numbers, this experiment illustrates several principles that are consistently undervalued in software engineering discourse:

1. Measure before you optimize. The difference between 2,500 and 5,200 users was achieved with two small caching changes. Not a rewrite. Not a new framework. Not a migration to a different database. Not adding three more servers. Measure first. Optimize the actual bottleneck.

2. Caching is the highest-leverage optimization in most systems. Moving data closer to the user — whether that’s in-process memory, a reverse proxy cache, a CDN, or a client-side service worker — eliminates entire layers of computation for repeated requests. The Nginx cache took raw throughput from 280 req/s to over 10,000 req/s. That’s a 35x improvement from a configuration change.

3. The “it depends” answer is real but not useless. The honest answer to “how many users can my server handle?” is “it depends on what the users are doing.” But that doesn’t mean you can’t get a concrete number for your specific workload. Model your users. Run the test. Get your number.

4. Simple architectures are not naive architectures. A single server running Nginx, an app server, and a database is not a “toy setup” that only amateurs use. It’s the correct architecture for the vast majority of applications at the vast majority of traffic levels. Complexity should be added incrementally, driven by measured need, not by architectural fashion.

5. Tail latency breaks first. At 2,000 users, the median experience was still excellent. The P95 and P99 were the first metrics to degrade. If you’re only monitoring average response time, you will miss the early warning signs of capacity exhaustion. Monitor percentiles.

Hardware and Cost Context

For completeness, here’s the cost breakdown of this entire experiment:

ResourceSpecificationDurationCost
Application server (DigitalOcean Basic Droplet)1 vCPU, 2 GB RAM, 50 GB SSD~1 week~$1.50
Load generation server (DigitalOcean)4 vCPU, 8 GB RAM~4 hours<$1.00
Total experiment cost<$2.50

The knowledge gained — exact capacity thresholds, failure modes, optimization impact — would cost thousands of dollars in engineering time if discovered during a production incident. Spending $2.50 to learn your system’s limits before your users do is one of the highest-ROI activities in operational engineering.

Final Assessment

A 1-core, 2 GB RAM server running Nginx, Node.js, and PostgreSQL on the same host can sustain 2,500 concurrent active users in its baseline configuration, handling approximately 232 requests per second with zero errors and sub-second P99 latency. With two targeted caching optimizations — an in-process Node.js cache and an Nginx proxy cache — that same hardware sustains 5,200 concurrent active users at over 500 requests per second.

The bottleneck is CPU. Always CPU, on a single-core machine. Not RAM. Not disk. Not network. The CPU saturates, the queue forms, and latency explodes non-linearly.

The fix, in most cases, isn’t a bigger server. It’s eliminating repeated work. Cache the response. Serve it closer to the user. Let Nginx handle what Nginx is extremely good at: serving cached bytes over TCP without engaging application logic.

The next time someone tells you that you need a Kubernetes cluster to serve a web application, point them at these numbers. Then ask them how many concurrent users they’re actually serving. The answer, for 95% of applications on the internet, is well within the capacity of a single $6 VPS.


Methodology note: All tests were conducted on DigitalOcean infrastructure in the NYC data center. The application was purpose-built for this experiment. Results are specific to the described workload pattern (social media feed with 90% reads, 10% writes, 3–15 second think times). Different workload characteristics — heavier writes, larger payloads, more complex queries, compute-intensive endpoints — will produce different capacity numbers. Test your own system.


Buy me A Coffee!

Support The CyberSec Guru’s Mission

🔐 Fuel the cybersecurity crusade by buying me a coffee! Your contribution powers free tutorials, hands-on labs, and security resources.

Why your support matters:
  • Writeup Access: Get complete writeup access within 12 hours
  • Zero paywalls: Keep the main content 100% free for learners worldwide

Perks for one-time supporters:
☕️ $5: Shoutout in Buy Me a Coffee
🛡️ $8: Fast-track Access to Live Webinars
💻 $10: Vote on future tutorial topics + exclusive AMA access

“Your coffee keeps the servers running and the knowledge flowing in our fight against cybercrime.”☕ Support My Work

Buy Me a Coffee Button

If you like this post, then please share it:

Glossary

Discover more from The CyberSec Guru

Subscribe to get the latest posts sent to your email!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from The CyberSec Guru

Subscribe now to keep reading and get access to the full archive.

Continue reading