pr1v8.ca

Hardening three non-exit Tor relays on one VPS: offline keys, file integrity and unattended alerts

· 33 min read · torself-hostinghardeningdebian

I run three Tor relays on one rented virtual server with 4 vCPUs, 8 GB of RAM and one public IPv4 address per relay. They are guard and middle relays, never exits. Relay Search measures about 30 MB/s of capacity between them, and they mostly run without me.

This guide covers how they are built and hardened. The identity keys live off the server, file changes are checked every day, memory is sized so the relays stay in the consensus, and alerts only arrive when a person is needed. It also covers the parts that failed first. It is for someone with a spare VPS who wants to give bandwidth to Tor without it turning into a second job.

Relays are public by design. Every relay’s address, flags, bandwidth and family are in the hourly consensus, so you can check what this post says against the real thing:

Relay Fingerprint (opens Relay Search)
BungeeNet1 EA560CB1EE1619B5E124E211F83D87D7C105D2F6
BungeeNet2 5CAB947F5311E08D2F65E291DBA4C857AB841AE2
BungeeNet3 4A7E967EB2E126E6D87C29E80346873C59387B52

In the rest of the post they are relay1, relay2 and relay3 in that order. Relay2 is the interesting one: it became a busy directory cache, and most of what I learned came from it.

What you need:

The order, each step useful on its own:

  1. The host
  2. One tor process per address
  3. The torrc
  4. A family, and proof that it is yours
  5. Memory decides how many relays fit
  6. Two systemd limits that hurt
  7. Restarts cost flags
  8. Keep the identity key off the server
  9. Being told, not watching
  10. File integrity without the noise
  11. The rest of the hardening
  12. Smaller things that went wrong
  13. What is still open

One server, three relays VPS: 4 vCPU, 8 GB RAM, Debian 13 relay1 tor@default IPv4 + IPv6, port 443 queues up to 3 GB relay2 tor@relay2 IPv4, port 443 queues up to 3 GB relay3 tor@relay3 IPv4, port 443 queues up to 1.5 GB control on unix sockets, metrics on 127.0.0.1 only timers: watcher, self-heal, daily AIDE and dpkg --verify at home: master identity keys mint a 6-month signing key push alerts over HTTPS only when the state changes

1. The host#

Install tor from the Tor Project’s own repository rather than Debian’s, so security releases arrive within days. The Tor Project documents the repository setup.

Then let the box patch itself. unattended-upgrades only installs from origins you list, so add the Tor Project’s, and let it reboot when a kernel update needs one:

// /etc/apt/apt.conf.d/52tor-relay
Unattended-Upgrade::Origins-Pattern { "origin=TorProject"; };
Unattended-Upgrade::Automatic-Reboot "true";

Each tor upgrade restarts every relay and each kernel upgrade reboots the server. Both cost one flag for four days (section 7). I keep it on anyway: an internet-facing relay running an unpatched tor or kernel is the worse problem.

Add the extra addresses to the network config as /32 addresses so they survive a reboot. Two things caught me here. My firewall only allows SSH from a known address, so testing a new address on port 22 made it look unrouted when it was fine; test on the port the relay will use. And my provider’s gateway took a few minutes to start answering for a freshly added address, so a test straight after ip addr add failed and the same test later passed.

2. One tor process per address#

Debian’s tor package runs several independent instances. The default one, tor@default, reads /etc/tor/torrc. Each extra instance gets its own config, system user, data directory and unit:

sudo tor-instance-create relay2
sudo tor-instance-create relay3

For relay2 that makes /etc/tor/instances/relay2/torrc, the user _tor-relay2, the data directory /var/lib/tor-instances/relay2/ and the unit tor@relay2.

Separate processes, not one tor with several ORPorts, for three reasons. Each relay needs its own identity. Tor’s main loop is single-threaded, so one process per relay is how a relay host uses more than one core. And restarting one relay leaves the others alone.

A generator in the package starts every instance at boot. systemctl status shows them as enabled-runtime with no [Install] section; that is normal, don’t run systemctl enable on them.

To check a config before restarting, run the same --verify-config command the unit runs. systemctl cat tor@relay2 shows it in ExecStartPre. An instance’s torrc depends on a defaults file with @@NAME@@ placeholders that the unit fills in first, so running tor --verify-config on the instance torrc by hand fails with an error that looks like a broken config when it isn’t.

3. The torrc#

Relay2’s config, with documentation addresses and an example domain in place of mine:

Nickname        ExampleRelay2
ContactInfo     ciissversion:3 email:tor[at]example[dot]com url:example.com proof:dns-familyid-ed25519

ORPort              198.51.100.12:443
Address             198.51.100.12
OutboundBindAddress 198.51.100.12

SocksPort   0
ExitRelay   0
ExitPolicy  reject *:*

ControlSocket        /run/tor-instances/relay2/control GroupWritable RelaxDirModeCheck
CookieAuthentication 1
MetricsPort          127.0.0.1:9036
MetricsPortPolicy    accept 127.0.0.1

FamilyId  <family-id>
MyFamily  $<relay1-fingerprint>,$<relay3-fingerprint>

MaxMemInQueues          3 GB
RelayBandwidthRate      10 MBytes
RelayBandwidthBurst     12 MBytes
MaxAdvertisedBandwidth  10 MBytes

OfflineMasterKey    1
SigningKeyLifetime  6 months
Sandbox             1

Line by line, where it isn’t obvious:

The memory and bandwidth lines are section 5, and the key lines are section 8.

4. A family, and proof that it is yours#

All relays run by one operator must declare themselves a family, so clients never put two of them in the same circuit. On a single server this is the most important setting in the file. Since tor 0.4.9 it is done with a family key:

tor --keygen-family myfamily

That writes myfamily.secret_family_key and prints a FamilyId line. Copy the key file into each relay’s keys/ directory, owned by that instance’s user with mode 600, and put the FamilyId line in every torrc. When you add a relay later, copy the same file. A newly generated key is a different family.

I also keep MyFamily listing the other fingerprints, for clients that don’t understand family IDs yet. The descriptors carry both.

Adding relay3 meant changing MyFamily on two running relays. That can be set live over the control socket with SETCONF, so neither restarted; then the same line went into both torrc files so it survives the next restart. Relay Search should show every member under effective_family and nothing under alleged_family. A one-sided claim lands in alleged_family.

To prove the relays are yours (the Authenticated Relay Operator ID), the ContactInfo above ends in proof:dns-familyid-ed25519, and one DNS record covers the whole family:

we-run-this-tor-ed25519-family-id.example.com.  IN TXT  "<family-id>"

The value must match FamilyId exactly, and the zone must be signed with DNSSEC. The operator setup guide I followed only described the older per-relay RSA proofs; the ContactInfo specification at ciissversion:3 defines the family-ID ones, and when the two disagree the specification wins.

5. Memory decides how many relays fit#

A month in, relay2 dropped out of the consensus for about five hours. systemctl showed it as active the whole time, and nothing restarted it, because nothing had crashed.

How a busy relay fell out of the consensus 1. The uplink fills no bandwidth cap; the link pinned at 496 of about 500 Mbit/s 2. Queues grow faster than they drain relay2 wrote 18 times what it read: directory answers 3. Tor's own memory limit is reached MaxMemInQueues 2 GB; its OOM handler frees 1.4 GB, 3/4 buffers 4. The single main thread saturates above 100% of one core, killing circuits instead of serving 5. New connections are not accepted 4,096 wait in the kernel; the authorities cannot reach it Out of the consensus for about five hours, while systemd said active

Each step in that chain was measured, not guessed. Step 3 is where the memory went. When tor’s own out-of-memory handler fired, it reported roughly 0.5 GB of queued cells and 1.5 GB of connection buffers every time. About three quarters of the memory was socket buffers for some 9,000 connections, not relayed data.

That explains why the first fix only half worked. I capped bandwidth (RelayBandwidthRate, MaxAdvertisedBandwidth), which stopped the link from saturating and stopped the immediate outage. But relay2’s CPU didn’t move at all: 108% of a core before the cap and after. What fixed it for good was giving the queues room: MaxMemInQueues from 2 GB to 3 GB. Twelve days later the out-of-memory handler had fired zero times, the accept queue sat at zero, and relay2 was still running at 106% of a core. A relay can sit above 100% of one core for weeks and be fine, as long as it isn’t also fighting its own memory limit.

relay1 and relay2 on one server, measured the same day CPU, % of one core relay1 28% relay2 106% Directory requests per second relay1 1 relay2 431 Circuits created per second relay1 7 relay2 228 Peak memory since the last restart, MB relay1 1,523 relay2 4,159 Bars are scaled per row. Load follows requests, not bytes.

Directory requests are small and the answers are big, so a directory cache writes far more than it reads and costs CPU per request, not per byte. Size a relay by its request load, not its bandwidth. My first cap was sized from megabits per core, and it was wrong.

How I size memory now:

Today the three relays sit at about 1.0 GB resident each, 3.1 GB of 7.76 GB is in use, and swap is empty.

6. Two systemd limits that hurt#

Both of these looked like sensible guard rails, and both made things worse.

CPUQuota throttles bursts. I capped each relay at CPUQuota=150%. One evening relay1 started logging “Your computer is too slow to handle this many circuit creation requests”, dropped a third of its legacy circuit handshakes (114,608 of 342,623) and fell out of the consensus. The server had cores to spare. CFS bandwidth control counts per 100 ms period, and tor’s worker threads burst inside a single period. After raising the quota to 300%, relay1 used 6.5% of its allowance and was still throttled 17 times in 90 seconds. Any finite quota does this under load. CPUWeight is the safe alternative: it only matters when cores are contended, and it never throttles against idle ones.

MemoryHigh reclaims instead of protecting. MemoryHigh is a soft limit: above it the kernel reclaims the cgroup’s memory constantly. Relay1 logged 878,673 such events while sitting just under its limit, spending CPU without ever being in danger. MemoryMax is a hard wall that does nothing until it is hit. Use that, set well above the measured peak, so a runaway relay can’t take the whole server and nothing happens otherwise.

The drop-ins I use now, one pair per relay. systemctl edit --drop-in=zz-cpu tor@relay2 opens the first one in the right directory:

# zz-cpu.conf
[Service]
CPUQuota=
CPUWeight=100
TasksMax=4096
# memory.conf
[Service]
MemoryHigh=infinity
MemoryMax=5G

The empty CPUQuota= clears any quota set by an earlier file. Drop-ins apply in lexicographic order and the last one wins, hence the zz- prefix.

Cgroup limits don’t need a tor restart: systemctl set-property --runtime changes them on a live unit. The trap is the word runtime. It writes to /run, which is gone at the next reboot. My original quota existed only there, so the next automatic reboot would have silently changed the limits. Put the values you want to keep in /etc, and check that nothing important lives only in /run:

systemctl show tor@relay2 -p DropInPaths

7. Restarts cost flags#

Flags decide how much traffic a relay gets, and most of them survive a short restart. Only one doesn’t:

Flag What the directory authorities look at After a one-minute restart
Running they reached it in the last 45 minutes usually nothing
Guard fraction of each day it was up, plus how long it has been known negligible
Stable a decaying average of how long each uptime lasted a small dip
HSDir at least 96 hours of continuous uptime gone for four days

Turning off automatic reboots to protect Guard looks sensible and isn’t: a one-minute reboot is 0.07% of a day. HSDir comes and goes with every tor or kernel update, roughly monthly, and that’s the price of patching.

Three habits keep restarts rare:

8. Keep the identity key off the server#

A relay’s permanent identity is an ed25519 master key. Normally it sits in the data directory, so anyone who copies the disk can impersonate the relay forever. With OfflineMasterKey 1, tor never loads the master key. It runs on a signing key the master key has certified for a limited time, here six months, and the master key lives somewhere else.

Moving an existing relay over, once:

  1. Copy the whole keys/ directory off the server and keep it somewhere safe. That copy is now the only thing that can renew the relay.
  2. On a trusted machine, mint a six-month signing key from that copy with the renewal command below, and put the two signing files in the relay’s keys/ directory.
  3. Add OfflineMasterKey 1 and SigningKeyLifetime 6 months to the torrc and restart the relay once.
  4. Delete ed25519_master_id_secret_key from the server and check the fingerprint didn’t change.

I did the one restart while HSDir was already missing after an unplanned reboot, so it cost nothing extra.

Renewal happens on a trusted machine at home, not on the server:

work=$(mktemp -d -p /dev/shm); mkdir -m 700 "$work/keys"
cp ed25519_master_id_secret_key ed25519_master_id_public_key "$work/keys/"
tor -f /dev/null --DataDirectory "$work" --SigningKeyLifetime "6 months" --keygen
# copy $work/keys/ed25519_signing_cert and ed25519_signing_secret_key to the relay's
# keys/ directory (owner = that instance's user, mode 600), then:
shred -u "$work"/keys/*; rm -rf "$work"

Two details in there are easy to get wrong:

My script for this pulls the key backup from my password manager into memory, checks that each master public key matches the one on the relay before minting anything, copies only the two signing files, and shreds the masters when it exits. It costs one unlock and one hardware-key touch.

Tor checks its keys on a timer and should pick up a newer signing certificate from disk by itself, so renewal shouldn’t need a restart. My first real renewal is months away, so I haven’t seen that happen yet. The script has passed a full dry run against all three relays.

The expiry date is inside the certificate file, after a 32-byte text header, as four big-endian bytes of hours since 1970:

h=$(od -An -tu4 --endian=big -j34 -N4 /var/lib/tor/keys/ed25519_signing_cert | tr -d ' ')
date -u -d @$((h*3600))

My watcher alerts weeks before expiry and again close to it. If renewal is missed the relays stop working, but nobody can take over their identity.

What this doesn’t cover: the legacy RSA identity key and the family key stay on the server. The RSA key can’t be moved offline the same way. The family key could be, but that is more work.

9. Being told, not watching#

A timer runs a check script every few minutes. It pushes a notification only when the overall state changes, to problem or back to healthy, so a long outage is one message, not a hundred. It sends over plain HTTPS rather than through Tor, so alerts still arrive when Tor is the thing that’s broken. The state only advances after a confirmed send, so an outage at the push service delays an alert rather than losing it.

curl -fsS --data-urlencode "token=<app-token>" --data-urlencode "user=<user-key>" \
     --data-urlencode "title=relay problem" --data-urlencode "message=$msg" \
     --data-urlencode "priority=1" https://api.pushover.net/1/messages.json

What it checks, and why each one is there:

curl -s http://127.0.0.1:9036/metrics | grep '^tor_relay_flag' | grep ' 1$'

Every flag is listed with 0 or 1, so leave out the filter and you get nonsense. - Accepted connections: this is the check that would have caught the outage in section 5. ss -ltn shows the listen backlog of each ORPort in its first column; a number that stays high means tor is up but not accepting. - Lost flags: it remembers which flags each relay had and alerts when one goes. It stays quiet after a restart until HSDir has had time to come back by itself. - Counters as differences: tor’s overload counters (dropped handshakes, OOM bytes, TCP port exhaustion) only ever go up. Compare each to its previous value. Alert on the absolute number and one bad evening keeps the alert on forever. - Exit tripwire: it reads ExitRelay from both the torrc and the running config over the control socket, because a relay’s running config can differ from its file. - Signing certificate expiry, from section 8.

Two more pieces: a boot service sends a quiet notification on every boot, so an unexpected reboot never goes unnoticed, and a separate follower copies tor’s notices to a rotated file. The journal here is kept in RAM only, and without the copy an overnight reboot would erase the evidence of whatever caused it. Tor’s SafeLogging is on by default, so client addresses never reach the log.

Outside the server, the Tor Project’s own relay alert service, Tor Weather, emails me if a relay goes offline. That covers the one case a watcher on the server can’t report: the server itself being down.

10. File integrity without the noise#

I wanted to know if anyone changed files on the server, so I installed AIDE, a file-integrity checker: it records hashes and metadata in a database and reports any difference each day. With Debian’s default rules it tracked over 41,000 files, including every kernel module. The first few alerts were my own changes and cache files, which I pruned one by one. Then a kernel update arrived, the server rebooted into it, and for three nights I woke up to a priority alert listing five thousand added and five thousand removed files. All of it was the upgrade. An alert that fires after every upgrade gets ignored, and then a real one gets ignored too.

What checks what AIDE: hashes and metadata, daily /etc, relay identity and family keys own scripts, login files, SSH keys, dpkg checksum lists dpkg --verify, daily every file a package installed (/usr, /boot) against the checksums its package shipped Not watched, on purpose files that change on their own: runtime state, caches, logs, rotating keys AIDE database, entries before 41,145 after 3,034

So the work is split:

dpkg --verify | grep -v ' c /' # drop config files; those are AIDE's job

On my server that prints a single line, a directory the AppArmor package expects that doesn’t exist here, so the daily check carries an allowlist for it. - Files that change on their own, such as runtime state, caches, logs and tor’s statistics, are left out of AIDE with prune rules (lines starting with !) in a file under /etc/aide/aide.conf.d/.

The database went from 41,145 entries to about 3,000. Getting there taught me four things about AIDE that the defaults don’t make obvious:

To prove the scope, I created a file in each kind of place and ran a check. New files in /usr/local/sbin and /etc/tor were reported, new files in /usr/bin and /boot were not. Then I learned that creating and deleting a file changes the parent directory’s timestamps, so the next morning’s check reported my own test. Re-baseline after a test like that.

A root attacker on a running server can rewrite the AIDE database, so this mainly catches changes made while the server was off, or by something without root. I keep the database immutable (chattr +i) and store its hash off the server; after every deliberate re-baseline the stored hash gets updated.

11. The rest of the hardening#

Briefly, because none of this is unusual:

12. Smaller things that went wrong#

13. What is still open#