When Does a Rust Rewrite Actually Pay Off? A Decision Checklist

Home Blog When Does a Rust Rewrite Actually Pay Off? A Decision Checklist

"Should we rewrite it in Rust?" is one of the most expensive questions an engineering team can answer wrongly in either direction. Say yes too easily and you spend six months re-implementing features users already had. Say no out of habit and you keep paying for servers, incidents and latency that a focused rewrite would have removed. This article gives you a way to decide on evidence, plus a checklist you can score your own system against.

First: what Rust actually buys you

Strip away the hype and Rust gives you three concrete things:

  1. Predictable performance. No garbage collector, so no GC pauses. Compiled to native code with control over memory layout, so CPU-bound code typically runs at C/C++ speed.
  2. Memory safety without a runtime. Whole classes of bugs (use-after-free, data races, buffer overflows) are rejected at compile time.
  3. Small, efficient deployables. A single static binary with low memory use, which matters at scale and at the edge.

Everything else, like "it's modern" or "developers like it", is a bonus and not a business case.

When a rewrite pays off

1. You're CPU-bound on a hot path

If profiling shows most time is spent in your own code (parsing, serialization, pricing, matching, compression, image or signal processing) rather than waiting on a database or network, Rust can turn that path from a bottleneck into a non-issue. Typical candidates: market data handlers, risk calculations, rule engines, ETL transforms.

2. Your problem is tail latency, not average latency

Garbage-collected runtimes can have excellent averages and ugly p99s. If your SLO is about the slowest 1% of requests (trading, real-time bidding, game servers, anything interactive), removing GC pauses can matter more than raw speed. Discord's often-cited move of its "Read States" service from Go to Rust was driven by exactly this: periodic latency spikes from garbage collection, which went away after the rewrite.

3. Infrastructure cost is a real line item

At small scale, a 3× efficiency gain saves a few hundred euros a month and isn't worth a rewrite. At large scale it's different: Cloudflare reported that Pingora, its Rust-based proxy that replaced its NGINX setup, used about 70% less CPU and 67% less memory for the same traffic. If compute is one of your biggest costs, run the numbers.

4. Correctness failures are expensive

Parsers handling untrusted input, financial calculations, concurrent state machines: places where a memory bug or race condition becomes a security incident or a wrong trade. Rust's compiler catches categories of these before they ship.

5. You need to run where runtimes are a burden

Edge devices, embedded systems, WebAssembly in the browser, CLI tools you distribute to customers: one small binary with no runtime to install is a genuine advantage.

When it doesn't pay off

  • The system is I/O-bound. If requests spend 95% of their time waiting on Postgres or a third-party API, a faster language makes the remaining 5% faster. Fix the queries, add caching, batch the calls.
  • The product is still changing weekly. Rust rewards knowing what you're building. Before product-market fit, iteration speed beats runtime speed.
  • Nobody will own it. If the team that maintains the system won't learn Rust, you're creating an island. Budget for training or a partner, or don't do it.
  • The real problem is architecture. A chatty microservice design, N+1 queries or a missing queue will be just as slow in Rust.
  • "Rewrite everything" is the plan. Big-bang rewrites of working systems are where projects go to die, regardless of language.

The approach that works: rewrite the hot path, not the system

In practice the best return comes from replacing the 5–10% of the code where the time goes, and leaving the rest alone:

  1. Profile production-like load. Flame graphs, not intuition. Find where CPU time and allocations actually go.
  2. Put a boundary around the hot path. Define its inputs and outputs precisely. This step alone often reveals simplifications.
  3. Rewrite that component in Rust and call it from the existing code. Rust integrates well with other stacks: PyO3 for Python, napi-rs for Node.js, a C ABI with P/Invoke for C#/.NET, or a separate service over gRPC/HTTP.
  4. Prove equivalence. Run old and new side by side on recorded production inputs and diff the outputs. Property-based tests help here.
  5. Measure, then decide on the next piece. Gradual replacement (the "strangler fig" pattern) lets you stop as soon as the remaining gains aren't worth it.

This keeps risk low, gives you real numbers within weeks instead of quarters, and often delivers most of the benefit of a full rewrite.

The checklist

Score each statement 0 (not true), 1 (partly) or 2 (clearly true) for the system you're considering:

# Statement Score
1 Profiling shows we're CPU-bound in our own code, not waiting on I/O.
2 Tail latency (p99/p999) or latency jitter is a business problem.
3 Compute or memory cost for this component is significant and growing.
4 Correctness or memory-safety failures here are costly (money, security, trust).
5 The component's behaviour is well understood and fairly stable.
6 We can isolate it behind a clear interface.
7 We have, or can get, people to build and maintain Rust code.
8 We can test the new version against real production inputs.

How to read your score:

  • 0–6: Not now. Look at architecture, queries and caching first.
  • 7–11: Run a time-boxed proof of concept on the single hottest path and measure.
  • 12–16: A targeted rewrite is very likely worth it. Plan it incrementally.

Statements 1–4 are the why; 5–8 are the can we. A high "why" with a low "can we" means: fix the prerequisites first (interfaces, tests, people) before touching the code.

What a proof of concept should answer

Keep it to two to four weeks and make it answer three questions with numbers: how much faster or cheaper is the hot path under realistic load, how hard was it to integrate with the existing system, and what will it cost to maintain. If the answers are good, you have a business case. If not, you've spent a few weeks instead of a few quarters.

We help teams with exactly this: profiling, isolating the hot path, building the Rust component and integrating it. See Rust development, legacy modernization and proof of concept development, or start with an independent build vs. buy evaluation.