Hardening three non-exit Tor relays on one VPS: offline keys, file integrity and unattended alerts
I run three Tor relays on one rented virtual server with 4 vCPUs, 8 GB of RAM and one public IPv4 address per relay. They are guard and middle relays, never exits. Relay Search measures about 30 MB/s of capacity between them, and they mostly run without me.
This guide covers how they are built and hardened. The identity keys live off the server, file changes are checked every day, memory is sized so the relays stay in the consensus, and alerts only arrive when a person is needed. It also covers the parts that failed first. It is for someone with a spare VPS who wants to give bandwidth to Tor without it turning into a second job.
Relays are public by design. Every relay’s address, flags, bandwidth and family are in the hourly consensus, so you can check what this post says against the real thing:
| Relay | Fingerprint (opens Relay Search) |
|---|---|
| BungeeNet1 | EA560CB1EE1619B5E124E211F83D87D7C105D2F6 |
| BungeeNet2 | 5CAB947F5311E08D2F65E291DBA4C857AB841AE2 |
| BungeeNet3 | 4A7E967EB2E126E6D87C29E80346873C59387B52 |
In the rest of the post they are relay1, relay2 and relay3 in that order. Relay2 is the interesting one: it became a busy directory cache, and most of what I learned came from it.
What you need:
- a VPS whose provider allows Tor relays (read the terms; mine allow non-exit relays)
- one public IPv4 address per relay; extra IPv4 addresses are cheap at most VPS hosts
- about 3 GB of RAM for each busy relay; RAM, not CPU or bandwidth, decides how many fit
- Debian 13. Any systemd distribution can do this, but the multi-instance tooling below is Debian’s
- a push-notification service with an HTTP API for alerts (I use Pushover)
- somewhere off the server to keep identity keys, such as a password manager
The order, each step useful on its own:
- The host
- One tor process per address
- The torrc
- A family, and proof that it is yours
- Memory decides how many relays fit
- Two systemd limits that hurt
- Restarts cost flags
- Keep the identity key off the server
- Being told, not watching
- File integrity without the noise
- The rest of the hardening
- Smaller things that went wrong
- What is still open
1. The host#
Install tor from the Tor Project’s own repository rather than Debian’s, so security releases arrive within days. The Tor Project documents the repository setup.
Then let the box patch itself. unattended-upgrades only installs from origins you list,
so add the Tor Project’s, and let it reboot when a kernel update needs one:
// /etc/apt/apt.conf.d/52tor-relay
Unattended-Upgrade::Origins-Pattern { "origin=TorProject"; };
Unattended-Upgrade::Automatic-Reboot "true";
Each tor upgrade restarts every relay and each kernel upgrade reboots the server. Both cost one flag for four days (section 7). I keep it on anyway: an internet-facing relay running an unpatched tor or kernel is the worse problem.
Add the extra addresses to the network config as /32 addresses so they survive a reboot.
Two things caught me here. My firewall only allows SSH from a known address, so testing
a new address on port 22 made it look unrouted when it was fine; test on the port the
relay will use. And my provider’s gateway took a few minutes to start answering for a
freshly added address, so a test straight after ip addr add failed and the same test
later passed.
2. One tor process per address#
Debian’s tor package runs several independent instances. The default one, tor@default,
reads /etc/tor/torrc. Each extra instance gets its own config, system user, data
directory and unit:
sudo tor-instance-create relay2
sudo tor-instance-create relay3
For relay2 that makes /etc/tor/instances/relay2/torrc, the user _tor-relay2, the data
directory /var/lib/tor-instances/relay2/ and the unit tor@relay2.
Separate processes, not one tor with several ORPorts, for three reasons. Each relay needs its own identity. Tor’s main loop is single-threaded, so one process per relay is how a relay host uses more than one core. And restarting one relay leaves the others alone.
A generator in the package starts every instance at boot. systemctl status shows them
as enabled-runtime with no [Install] section; that is normal, don’t run
systemctl enable on them.
To check a config before restarting, run the same --verify-config command the unit
runs. systemctl cat tor@relay2 shows it in ExecStartPre. An instance’s torrc depends on
a defaults file with @@NAME@@ placeholders that the unit fills in first, so running
tor --verify-config on the instance torrc by hand fails with an error that looks like a
broken config when it isn’t.
3. The torrc#
Relay2’s config, with documentation addresses and an example domain in place of mine:
Nickname ExampleRelay2
ContactInfo ciissversion:3 email:tor[at]example[dot]com url:example.com proof:dns-familyid-ed25519
ORPort 198.51.100.12:443
Address 198.51.100.12
OutboundBindAddress 198.51.100.12
SocksPort 0
ExitRelay 0
ExitPolicy reject *:*
ControlSocket /run/tor-instances/relay2/control GroupWritable RelaxDirModeCheck
CookieAuthentication 1
MetricsPort 127.0.0.1:9036
MetricsPortPolicy accept 127.0.0.1
FamilyId <family-id>
MyFamily $<relay1-fingerprint>,$<relay3-fingerprint>
MaxMemInQueues 3 GB
RelayBandwidthRate 10 MBytes
RelayBandwidthBurst 12 MBytes
MaxAdvertisedBandwidth 10 MBytes
OfflineMasterKey 1
SigningKeyLifetime 6 months
Sandbox 1
Line by line, where it isn’t obvious:
ORPort 443: clients behind strict firewalls can usually still reach 443. There is no DirPort: it is deprecated, and directory requests come in over the ORPort.AddressandOutboundBindAddresspin the instance to its own address in both directions. WithoutOutboundBindAddress, every relay’s outgoing connections leave from the server’s primary address.SocksPort 0: Debian’s defaults open a SOCKS port; a relay doesn’t need one.ExitRelay 0andExitPolicy reject *:*: either one is enough to stay non-exit; I set both. Exit relays draw abuse complaints and attention to whoever runs them; a middle relay carries the same encrypted traffic without being anyone’s last hop.ControlSocketis a unix socket with cookie authentication, never a TCP control port.GroupWritablelets a local console user in the instance’s group read it.MetricsPortstays on localhost. The Tor Project asks operators not to publish per-relay statistics, and the watcher in section 9 reads it locally.DirCacheisn’t listed, so it stays at the default of 1. I triedDirCache 0on the third relay to save memory, and--verify-configsaid: “DirCache is disabled and we are configured as a relay. We will not become a Guard.” A relay that doesn’t serve directory information stays a middle relay forever.Sandbox 1turns on tor’s seccomp filter. The side effect is in section 7:systemctl reloadbecomes a restart.
The memory and bandwidth lines are section 5, and the key lines are section 8.
4. A family, and proof that it is yours#
All relays run by one operator must declare themselves a family, so clients never put two of them in the same circuit. On a single server this is the most important setting in the file. Since tor 0.4.9 it is done with a family key:
tor --keygen-family myfamily
That writes myfamily.secret_family_key and prints a FamilyId line. Copy the key file
into each relay’s keys/ directory, owned by that instance’s user with mode 600, and put
the FamilyId line in every torrc. When you add a relay later, copy the same file. A
newly generated key is a different family.
I also keep MyFamily listing the other fingerprints, for clients that don’t understand
family IDs yet. The descriptors carry both.
Adding relay3 meant changing MyFamily on two running relays. That can be set live over
the control socket with SETCONF, so neither restarted; then the same line went into both
torrc files so it survives the next restart. Relay Search should show every member under
effective_family and nothing under alleged_family. A one-sided claim lands in
alleged_family.
To prove the relays are yours (the Authenticated Relay Operator ID), the ContactInfo
above ends in proof:dns-familyid-ed25519, and one DNS record covers the whole family:
we-run-this-tor-ed25519-family-id.example.com. IN TXT "<family-id>"
The value must match FamilyId exactly, and the zone must be signed with DNSSEC. The
operator setup guide I followed only described the older per-relay RSA proofs; the
ContactInfo specification at ciissversion:3 defines the family-ID ones, and when the two
disagree the specification wins.
5. Memory decides how many relays fit#
A month in, relay2 dropped out of the consensus for about five hours. systemctl showed
it as active the whole time, and nothing restarted it, because nothing had crashed.
Each step in that chain was measured, not guessed. Step 3 is where the memory went. When tor’s own out-of-memory handler fired, it reported roughly 0.5 GB of queued cells and 1.5 GB of connection buffers every time. About three quarters of the memory was socket buffers for some 9,000 connections, not relayed data.
That explains why the first fix only half worked. I capped bandwidth
(RelayBandwidthRate, MaxAdvertisedBandwidth), which stopped the link from saturating
and stopped the immediate outage. But relay2’s CPU didn’t move at all: 108% of a core
before the cap and after. What fixed it for good was giving the queues room:
MaxMemInQueues from 2 GB to 3 GB. Twelve days later the out-of-memory handler had fired
zero times, the accept queue sat at zero, and relay2 was still running at 106% of a core.
A relay can sit above 100% of one core for weeks and be fine, as long as it isn’t also
fighting its own memory limit.
Directory requests are small and the answers are big, so a directory cache writes far more than it reads and costs CPU per request, not per byte. Size a relay by its request load, not its bandwidth. My first cap was sized from megabits per core, and it was wrong.
How I size memory now:
- Look at the peak, not the average. Relay2’s cgroup peaked at 4,159 MB, about 1.1 GB over its 3 GB queue limit. Relay1, with the same limit, has stayed under 2.5 GB.
- Keep the realistic case in RAM. The three queue limits (3, 3 and 1.5 GB) add up to
nearly all of the server’s RAM, and the peaks would not fit together. They haven’t
happened together: the realistic bad case is the busy relay at its peak and the other two at
their normal size, which fits.
MemoryMax(section 6) stops any one relay from taking the whole server, and swap covers the rest. - Use swap as a backstop, never as capacity. Tor queues that get paged out add
latency, the directory authorities read that as an unreachable relay, and you are back
at step 5 of the chain. I run
vm.swappiness=1, and the kernel’s count of pages swapped out has read zero at every check. - Encrypt the swap with a throwaway key. A relay’s memory holds circuit keys, and the
provider holds the disk. My swap is a file behind dm-crypt with a random key read from
/dev/urandomat every boot and never stored. I set it up from a small boot service rather than/etc/crypttab, because a crypttab failure can hold up the boot and this server is only reachable through the provider’s console when that happens.
Today the three relays sit at about 1.0 GB resident each, 3.1 GB of 7.76 GB is in use, and swap is empty.
6. Two systemd limits that hurt#
Both of these looked like sensible guard rails, and both made things worse.
CPUQuota throttles bursts. I capped each relay at CPUQuota=150%. One evening relay1
started logging “Your computer is too slow to handle this many circuit creation requests”,
dropped a third of its legacy circuit handshakes (114,608 of 342,623) and fell out of the
consensus. The server had cores to spare. CFS bandwidth control counts per 100 ms
period, and tor’s worker threads burst inside a single period. After raising the quota to 300%, relay1
used 6.5% of its allowance and was still throttled 17 times in 90 seconds. Any finite
quota does this under load. CPUWeight is the safe alternative: it only matters when cores
are contended, and it never throttles against idle ones.
MemoryHigh reclaims instead of protecting. MemoryHigh is a soft limit: above it the
kernel reclaims the cgroup’s memory constantly. Relay1 logged 878,673 such events while
sitting just under its limit, spending CPU without ever being in danger. MemoryMax is a
hard wall that does nothing until it is hit. Use that, set well above the measured
peak, so a runaway relay can’t take the whole server and nothing happens otherwise.
The drop-ins I use now, one pair per relay. systemctl edit --drop-in=zz-cpu tor@relay2
opens the first one in the right directory:
# zz-cpu.conf
[Service]
CPUQuota=
CPUWeight=100
TasksMax=4096
# memory.conf
[Service]
MemoryHigh=infinity
MemoryMax=5G
The empty CPUQuota= clears any quota set by an earlier file. Drop-ins apply in
lexicographic order and the last one wins, hence the zz- prefix.
Cgroup limits don’t need a tor restart: systemctl set-property --runtime changes them on
a live unit. The trap is the word runtime. It writes to /run, which is gone at the next
reboot. My original quota existed only there, so the next automatic reboot would have
silently changed the limits. Put the values you want to keep in /etc, and check that
nothing important lives only in /run:
systemctl show tor@relay2 -p DropInPaths
7. Restarts cost flags#
Flags decide how much traffic a relay gets, and most of them survive a short restart. Only one doesn’t:
| Flag | What the directory authorities look at | After a one-minute restart |
|---|---|---|
| Running | they reached it in the last 45 minutes | usually nothing |
| Guard | fraction of each day it was up, plus how long it has been known | negligible |
| Stable | a decaying average of how long each uptime lasted | a small dip |
| HSDir | at least 96 hours of continuous uptime | gone for four days |
Turning off automatic reboots to protect Guard looks sensible and isn’t: a one-minute reboot is 0.07% of a day. HSDir comes and goes with every tor or kernel update, roughly monthly, and that’s the price of patching.
Three habits keep restarts rare:
- Never
systemctl reloada sandboxed tor. WithSandbox 1, the reload signal makes tor exit and systemd starts it again. A reload is a restart. - Change what you can live. Bandwidth,
MaxMemInQueues,MyFamilyandDirCachecan all be set over the control socket withSETCONF(nyx, stem or a few lines of Python will do it). Then write the same values into the torrc. Runtime-only values disappear at the next restart; mine survived an automatic tor upgrade because they were in the file. - After an outage, wait. A relay that stopped accepting connections will miss the next consensus and come back an hour or two later on its own. Restarting it resets the reachability checks it needs to get back in.
8. Keep the identity key off the server#
A relay’s permanent identity is an ed25519 master key. Normally it sits in the data
directory, so anyone who copies the disk can impersonate the relay forever. With
OfflineMasterKey 1, tor never loads the master key. It runs on a signing key the master
key has certified for a limited time, here six months, and the master key lives somewhere
else.
Moving an existing relay over, once:
- Copy the whole
keys/directory off the server and keep it somewhere safe. That copy is now the only thing that can renew the relay. - On a trusted machine, mint a six-month signing key from that copy with the renewal
command below, and put the two signing files in the relay’s
keys/directory. - Add
OfflineMasterKey 1andSigningKeyLifetime 6 monthsto the torrc and restart the relay once. - Delete
ed25519_master_id_secret_keyfrom the server and check the fingerprint didn’t change.
I did the one restart while HSDir was already missing after an unplanned reboot, so it cost nothing extra.
Renewal happens on a trusted machine at home, not on the server:
work=$(mktemp -d -p /dev/shm); mkdir -m 700 "$work/keys"
cp ed25519_master_id_secret_key ed25519_master_id_public_key "$work/keys/"
tor -f /dev/null --DataDirectory "$work" --SigningKeyLifetime "6 months" --keygen
# copy $work/keys/ed25519_signing_cert and ed25519_signing_secret_key to the relay's
# keys/ directory (owner = that instance's user, mode 600), then:
shred -u "$work"/keys/*; rm -rf "$work"
Two details in there are easy to get wrong:
-f /dev/null: without it,--keygenreads the machine’s own/etc/tor/torrc. On my desktop that is a client config and keygen refused to parse it.- Never run
--keygenon an empty directory. Without a master key present it starts creating a new one and waits for a passphrase on the terminal, and a new master key is a new relay.
My script for this pulls the key backup from my password manager into memory, checks that each master public key matches the one on the relay before minting anything, copies only the two signing files, and shreds the masters when it exits. It costs one unlock and one hardware-key touch.
Tor checks its keys on a timer and should pick up a newer signing certificate from disk by itself, so renewal shouldn’t need a restart. My first real renewal is months away, so I haven’t seen that happen yet. The script has passed a full dry run against all three relays.
The expiry date is inside the certificate file, after a 32-byte text header, as four big-endian bytes of hours since 1970:
h=$(od -An -tu4 --endian=big -j34 -N4 /var/lib/tor/keys/ed25519_signing_cert | tr -d ' ')
date -u -d @$((h*3600))
My watcher alerts weeks before expiry and again close to it. If renewal is missed the relays stop working, but nobody can take over their identity.
What this doesn’t cover: the legacy RSA identity key and the family key stay on the server. The RSA key can’t be moved offline the same way. The family key could be, but that is more work.
9. Being told, not watching#
A timer runs a check script every few minutes. It pushes a notification only when the overall state changes, to problem or back to healthy, so a long outage is one message, not a hundred. It sends over plain HTTPS rather than through Tor, so alerts still arrive when Tor is the thing that’s broken. The state only advances after a confirmed send, so an outage at the push service delays an alert rather than losing it.
curl -fsS --data-urlencode "token=<app-token>" --data-urlencode "user=<user-key>" \
--data-urlencode "title=relay problem" --data-urlencode "message=$msg" \
--data-urlencode "priority=1" https://api.pushover.net/1/messages.json
What it checks, and why each one is there:
- Each unit is active and its control socket answers.
- Each relay is in the consensus, from the flags gauge on the local MetricsPort:
curl -s http://127.0.0.1:9036/metrics | grep '^tor_relay_flag' | grep ' 1$'
Every flag is listed with 0 or 1, so leave out the filter and you get nonsense.
- Accepted connections: this is the check that would have caught the outage in
section 5. ss -ltn shows the listen backlog of each ORPort in its first column; a
number that stays high means tor is up but not accepting.
- Lost flags: it remembers which flags each relay had and alerts when one goes. It
stays quiet after a restart until HSDir has had time to come back by itself.
- Counters as differences: tor’s overload counters (dropped handshakes, OOM bytes, TCP
port exhaustion) only ever go up. Compare each to its previous value. Alert on the
absolute number and one bad evening keeps the alert on forever.
- Exit tripwire: it reads ExitRelay from both the torrc and the running config over
the control socket, because a relay’s running config can differ from its file.
- Signing certificate expiry, from section 8.
Two more pieces: a boot service sends a quiet notification on every boot, so an
unexpected reboot never goes unnoticed, and a separate follower copies tor’s notices to a
rotated file. The journal here is kept in RAM only, and without the copy an overnight
reboot would erase the evidence of whatever caused it. Tor’s SafeLogging is on by
default, so client addresses never reach the log.
Outside the server, the Tor Project’s own relay alert service, Tor Weather, emails me if a relay goes offline. That covers the one case a watcher on the server can’t report: the server itself being down.
10. File integrity without the noise#
I wanted to know if anyone changed files on the server, so I installed AIDE, a file-integrity checker: it records hashes and metadata in a database and reports any difference each day. With Debian’s default rules it tracked over 41,000 files, including every kernel module. The first few alerts were my own changes and cache files, which I pruned one by one. Then a kernel update arrived, the server rebooted into it, and for three nights I woke up to a priority alert listing five thousand added and five thousand removed files. All of it was the upgrade. An alert that fires after every upgrade gets ignored, and then a real one gets ignored too.
So the work is split:
- AIDE watches what no package owns:
/etc, the relays’ identity and family keys, my own scripts in/usr/local, the login files and SSH keys. These almost never change, so any report means something. - AIDE also watches the checksum lists dpkg keeps in
/var/lib/dpkg/info, becausedpkg --verifytrusts them. My first version pruned all of/var/lib/dpkg, and then someone able to edit the disk could replace a binary and its checksum entry together, and neither check would notice. Upgrades rewrite those lists too; my alert script tells an upgrade apart from an unexplained change instead of excluding them. dpkg --verifywatches what packages own. It compares every installed file with the checksums its package shipped, which is what you want from a “has a binary been swapped” check, and package upgrades don’t trip it:
dpkg --verify | grep -v ' c /' # drop config files; those are AIDE's job
On my server that prints a single line, a directory the AppArmor package expects that
doesn’t exist here, so the daily check carries an allowlist for it.
- Files that change on their own, such as runtime state, caches, logs and tor’s
statistics, are left out of AIDE with prune rules (lines starting with !) in a file
under /etc/aide/aide.conf.d/.
The database went from 41,145 entries to about 3,000. Getting there taught me four things about AIDE that the defaults don’t make obvious:
- A trailing
$changes the match. AIDE rules are regular expressions anchored at the start only.!/var/lib/tor/stats$matches the directory itself and nothing in it, so the statistics files inside kept showing up every day.stats/matches what is inside. - Order matters across files. AIDE first finds the deepest directory node a rule can
hang from, then uses the first matching rule in that node. A narrow rule of mine in a
later file lost to a broad Debian rule in an earlier one. My exceptions for routine
churn live in a file named to sort before Debian’s (
69_before70_). - Some keys rotate. Tor replaces its medium-term onion keys every 28 days and keeps
the previous one as
.old. That produced a priority alert a month after setup. Watch the identity keys and excludesecret_onion_key*. - Directory timestamps move on every upgrade. Packages write temporary files into
/etcand its subdirectories while installing, so directory modification times change even when nothing in them did. For directories under/etcI check ownership and permissions but not timestamps; a new file still shows up as its own entry.
To prove the scope, I created a file in each kind of place and ran a check. New files in
/usr/local/sbin and /etc/tor were reported, new files in /usr/bin and /boot were
not. Then I learned that creating and deleting a file changes the parent directory’s
timestamps, so the next morning’s check reported my own test. Re-baseline after a test
like that.
A root attacker on a running server can rewrite the AIDE database, so this mainly catches
changes made while the server was off, or by something without root. I keep the database
immutable (chattr +i) and store its hash off the server; after every deliberate
re-baseline the stored hash gets updated.
11. The rest of the hardening#
Briefly, because none of this is unusual:
- Firewall: nftables, default drop. Only the ORPorts are open to the world; SSH is not. There is no per-source connection limit on the ORPort, because a guard relay legitimately sees many clients behind one carrier-grade NAT address. A broad global rate limit, SYN cookies and tor’s own DoS defences, which adjust from the consensus, handle floods.
- Network sysctls:
fqwith BBR, 16 MB socket buffers, a wide local port range. Tor’s overload documentation suggests tuning these; mine already exceeded what it lists, and none of my incidents were network-stack problems. - AppArmor: Debian ships a
system_torprofile, applied throughAppArmorProfile=system_torin a unit drop-in. One trap: the stocktor@.servicesetsReadOnlyDirectories=/, which also makes/proc/self/attrread-only, and that is where systemd writes the profile change. My relays hadProtectKernelTunables=yesin their hardening drop-in, which happens to leave that path writable. A later, less hardened instance could not be confined at all. Check which profile a process runs under withcat /proc/<pid>/attr/current. - Kernel: module loading is switched off about 90 seconds after boot, so anything that
needs a module (the encrypted swap needs
dm_cryptandloop) has to start before that. Unprivileged user namespaces are off. The boot line carries the usual hardening options:init_on_alloc=1 randomize_kstack_offset=1 slab_nomerge vsyscall=none lockdown=integrity. I left outinit_on_free=1because relay2 has no CPU to spare. - SSH: keys only, from known addresses. There is also a second way in that doesn’t depend on my address: an onion service, on its own small tor instance, with client authorization. A separate instance means changing it never restarts a relay.
12. Smaller things that went wrong#
- A console couldn’t read the new relay’s control socket (“Permission denied”, errno
13). The console user wasn’t in
_tor-relay3. Add the user to each instance’s group and log in again. - Testing as root hides permission bugs. Root ignores group permissions, so a check I
ran through
sudopassed while the real console failed. Test as the user who will use the thing. pkill -f patternkilled my own session, because the pattern was also on my SSH command line. Find the PID and kill that.- Relay Search lags. Onionoo, behind Relay Search, ran one to two hours behind the consensus; don’t judge a change made in the last two hours there.
- The overload marker lasts 72 hours. Relay Search keeps a relay marked as overloaded for three days after it recovers. That’s a free check: if the marker clears on schedule, the fix held.
13. What is still open#
- HSDir resets with every tor or kernel update. Accepted.
- Relay2 has no CPU headroom. It has been stable at about 106% of one core for weeks,
but more directory demand would push it over. Lowering
MaxAdvertisedBandwidthwas supposed to shed directory load, but twelve days later relay2 still answered about 430 directory requests a second.DirCache 0would shed it, at the cost of the Guard and HSDir flags, which I’m not willing to pay. - The legacy RSA identity key and the family key are still on the server.
- The configuration lives on the server. Rebuilding from scratch means repeating this post by hand. Keeping it as code is the next job.