Files
keydr/docs/BUILDING.md
T
thallada f8d3bb5b2b docs(build): record the thin-binary layout and its measured effect
Replace the 'known remaining inefficiency' section with the layout that
now exists, plus the measured before/after (clean release build: 2m14s
-> 1m46s wall, keydr units 61.0s -> 32.4s, peak ~2.0 GB unchanged).

Add a maintenance note: new modules go in src/lib.rs, not src/main.rs.
Re-declaring them in the binary would silently restore the double
compile this refactor removed.
2026-08-07 01:08:44 +00:00

214 lines
7.7 KiB
Markdown

# Building keydr
## TL;DR
```bash
cargo build
cargo test
cargo build --release
```
Plain cargo. No wrapper, no memory ceiling, no reduced optimisation.
A clean release build peaks at **~2.0 GB** and takes **1m46s** on a
4-core / 7.6 GB VM; an incremental rebuild after touching one file is a
few seconds.
If that's not what you're seeing, read on.
---
## The build used to be unusable — here's what was actually wrong
Release builds consumed 5+ GB and never finished; the VM thrashed so
badly that SSH stopped responding and needed a console reboot.
The obvious explanations were all wrong:
- **Not too many parallel jobs.** The peak came from a *single* rustc
process. A single process's memory is unaffected by `--jobs`.
- **Not insufficient RAM.** 8 GB is fine for a project this size.
- **Not LTO** (though LTO made it worse).
**The cause was a codegen bug in `rust-i18n` v3.** Its macro emitted one
`HashMap::from([...])` construction *and* one `add_translations()` call
**per translation string** — all inside a single function body. With 21
locales and ~9,000 keys that's 8,652 separate HashMap constructions in
one function. LLVM's optimiser scales superlinearly on function size, so
it consumed ~4.5 GB trying to optimise that one initialiser.
Verified by expanding the macro:
| | rust-i18n v3.1.5 | rust-i18n v4.2.1 |
|---|---|---|
| `HashMap::from` constructions | 8,652 | **0** |
| `add_translations` calls | 8,652 | **21** (one per locale) |
| Expanded lines (lib target) | 93,026 | 73,444 |
Upstream fixed this in **v4.0.0**. Reproduce the check yourself:
```bash
cargo install cargo-expand
cargo expand --lib | grep -c 'HashMap::from'
```
### The fix
```bash
cargo add rust-i18n@4
```
That's it. Measured on this box, same machine, clean builds:
| Configuration | Peak | Time |
|---|---|---|
| v3, `opt-level=3`, thin LTO | 5.2 GB | stalled >18 min, never finished |
| v3, `opt-level=3`, no LTO | 4.5 GB | stalled >12 min, never finished |
| v3, `opt-level=1`, no LTO (workaround) | 1.8 GB | 15m30s |
| **v4, `opt-level=3`, thin LTO** | **2.08 GB** | **3m25s** |
| **v4, stock cargo defaults** | **1.92 GB** | **2m17s** |
The earlier workaround — dropping `opt-level` to 1 and disabling LTO —
has been **removed**. It was treating a symptom.
---
## What's still in `.cargo/config.toml`, and why
Nothing there reduces the optimisation of shipped code.
- **`jobs = 3`** — uses 3 of 4 cores so the box stays interactive while
compiling. Purely about responsiveness; delete it if you don't care.
- **`lld` linker** — ships with the Rust toolchain, no install needed.
GNU ld is single-threaded and holds the whole link graph in memory.
- **`debug = 1` for this crate, `debug = 0` for dependencies** (dev/test
profiles only). Debug info is the largest contributor to `target/` size
and link time. Level 1 keeps line numbers, so backtraces still work.
`[profile.release]` is deliberately absent — cargo's defaults are right.
---
## Project layout
All application code lives in the **library** (`src/lib.rs``src/run.rs`
and the module tree). `src/main.rs` is a 10-line entry point that calls
`keydr::run()`.
This matters for build time. `main.rs` used to re-declare the whole module
tree (`mod app; mod config; ...`), which made the binary a *second
independent crate*: every module — and the entire rust-i18n translation
table — was compiled twice, and 281 unit tests ran twice.
Measured effect of consolidating it, clean release build:
```
before after
keydr (bin) 33.7s 0.3s
keydr (lib) 26.0s 31.0s
keydr total 61.0s 32.4s (-28.6s)
total unit-seconds 292.2s 264.5s
wall clock 2m14s 1m46s (-28s, -21%)
```
Test coverage is unchanged — 341 unique tests before and after. The
execution count dropping from 622 to 340 is the duplicate run disappearing.
**Keep it this way:** if you add a module, declare it in `src/lib.rs`, not
`src/main.rs`. Re-declaring modules in the binary would silently restore
the double compile.
The remaining ~79% of build time is dependency compilation (`reqwest`
12.7s, `clap_builder` 11.4s, `toml_edit` 9.3s, `tokio` 8.0s...), which
incremental builds skip entirely.
---
## If a build ever does go wild again
The machine-level protections below need no per-project changes.
### Keep SSH alive (recommended, needs root once)
```bash
# Reserve memory for the SSH daemon so it can't be swapped out entirely.
sudo systemctl edit ssh # [Service] / MemoryMin=128M
```
**Do not bother with `systemd-oomd` on this box** — it is a separate
package whose only reverse-dependencies are `ubuntu-desktop*`, so Ubuntu
Server never installs it (`systemctl is-enabled systemd-oomd``not-found`
here). Use **earlyoom** instead, which is in `universe` and kills only the
single highest-scoring process rather than a whole cgroup:
```bash
sudo apt install earlyoom
# /etc/default/earlyoom
EARLYOOM_ARGS="-m 8 -s 5 -r 60 \
--avoid '(^|/)(systemd|sshd|mosh-server|tmux.*|bash|fish)$' \
--prefer '(^|/)(rustc|cargo|ld|lld|collect2)$'"
```
`--avoid` protects your shell; `--prefer` points it at the compiler. Fedora
enabled earlyoom by default for exactly this "system becomes completely
unresponsive, user has no choice but to force power off" scenario.
### Cap every build at once, forever (needs root once)
This is the global knob — no wrapper script, no per-project config:
```bash
sudo systemctl set-property user-1000.slice MemoryHigh=5G MemoryMax=6500M
```
Applies to every shell and every build you start, persists across reboots.
Because `sshd` itself lives in `system.slice`, capping the user slice can
never lock you out of a new login. (Ubuntu already ships `TasksMax=33%`
this way via `/usr/lib/systemd/system/user-.slice.d/`, so it's a
distro-blessed pattern.)
Note `MemoryHigh` at the *slice* level is reasonable — it throttles a
sprawling session gradually. Do not put it on a single short-lived build
scope, where it causes the reclaim-stall described below.
### Cap one build ad-hoc (no root)
`memory` is delegated to the user slice on this box, so you can confine a
single command without any wrapper script:
```bash
systemd-run --user --scope -p MemoryMax=4G -p MemorySwapMax=0 \
cargo build --release
```
`MemorySwapMax=0` is the important half. The lockup was never really an
OOM — it was *thrashing*. With swap available the kernel pages `sshd` out
to feed the build and the box goes catatonic while `oom_kill` stays at 0.
Denying the build swap turns a slow total failure into a fast contained one.
**Do not add `MemoryHigh` to a single build scope.** It sounds safer but
traps the process in continuous reclaim so it grinds forever instead of
dying. Measured against an identical 256 MB ceiling: `MemoryMax` alone →
clean kill in seconds; `MemoryMax` + `MemoryHigh` → still spinning after 45
seconds.
### Diagnosing which crate is expensive
```bash
cargo build --release --timings # HTML report in target/cargo-timings/
cargo install cargo-llvm-lines
cargo llvm-lines --release | head -30 # which functions generate the most IR
cargo expand --lib | wc -l # how much code a macro really emits
```
`cargo llvm-lines` and `cargo expand` are what actually found this bug.
Reach for them before touching `opt-level`.
## Upstream context
Cargo has no memory-aware job scheduling; it schedules on core count only.
That's a known gap ([rust-lang/cargo#12912](https://github.com/rust-lang/cargo/issues/12912)),
and maintainers have said they'd rather delegate resource limits to the OS
(cgroups) than build it into cargo. So the systemd approach above isn't a
hack — it's the sanctioned answer. But in this case none of it was needed:
the real fix was a dependency upgrade.