Monitoring
Watching Nixt Server — Prometheus metrics, health and readiness endpoints, logs, the doctor command and its every check, and what to alert on.
Nixt Server gives you four ways to see how it is doing: Prometheus metrics, health and readiness endpoints for your load balancer or orchestrator, logs, and doctor, which checks the things that stop mail arriving or being believed.
The metrics and health endpoint
The endpoint is off until you give it an address:
[telemetry]
metrics_bind = "127.0.0.1:9187"
It serves plain HTTP with no authentication, so bind it to a loopback or private address only. Without metrics_bind, metrics are still counted but not served, and the log says no metrics endpoint configured; metrics are kept but not served.
| Path | Answer |
|---|---|
/metrics | Every metric, in the Prometheus text format. |
/healthz | 200 ok whenever the process is running. Use it for liveness. |
/readyz | 200 when every readiness check passes, 503 otherwise, with one line per check. Use it for readiness. |
| Anything else | 404 This endpoint serves /metrics, /healthz and /readyz. |
Readiness
/readyz has one check, store, sampled every 15 seconds:
| Body | Meaning |
|---|---|
store: ok | The store answered the last read. |
store: not ready (not checked yet) | The node started moments ago. |
store: not ready (<error>) | The store did not answer; the log says the store did not answer a health read. |
Metrics
No metric label ever carries an address, a mailbox or anything else belonging to one person.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
vsx_build_info | gauge | version, node | Always 1. Tells a dashboard which version each node runs. |
vsx_messages_accepted_total | counter | role: mx, submission, jmap | Messages accepted onto the queue. |
vsx_messages_accepted_bytes_total | counter | role | Bytes of those messages. |
vsx_messages_refused_total | counter | role, reason | Messages refused, counted as role="mx", reason="filter" for the filter’s refusals. |
vsx_authentications_total | counter | protocol: submission, imap, pop3, jmap, managesieve, dav; outcome: ok, failed, unavailable, refused | Sign-in attempts. unavailable is a sign-in that could not be decided because the directory it checks against could not be reached; refused is a right password or token the access rules kept out. |
vsx_sending_refusals_total | counter | policy: send-as, sending-limit, sending-ceiling, unproved-domain | Messages the organisation’s sending policies refused, each also in the audit log. |
vsx_sessions_closed_total | counter | listener: mx, submission, submissions, imap, imaps, pop3, pop3s, managesieve; reason: harvesting, network-refused, per-address, per-network, address-rate, network-rate | Connections closed for guessing at addresses, connections refused before their greeting because their network kept doing so, and connections over a listener’s ceilings. |
vsx_filter_stages_cut_total | counter | stage | Scoring stages the filter’s time budget cut before a message was accepted, and left for delivery to finish. |
vsx_filter_verdicts_total | counter | verdict: deliver, tag, quarantine, hold, reject, tempfail; stage: the filter stage that decided, or thresholds when the score did | Messages arriving on port 25, by what the filter decided and which stage decided it. |
vsx_outbound_recipients_total | counter | outcome: delivered, bounced, deferred; reply: 2xx, 4xx, 5xx, none; provider: google.com, outlook.com, yahoo.com, icloud.com or other | Each recipient of each delivery attempt, by what it came to and the kind of reply that decided it. Mail that could not be sent at all, such as to a domain with no mail server, is counted too. |
vsx_outbound_connections_total | counter | tls: verified, unverified, none; provider | Connections opened to other mail servers, by how they were encrypted. |
vsx_connections_open | gauge | listener: mx, submission, submissions, imap, imaps, pop3, pop3s, managesieve, https | Connections each listener has open now. https counts JMAP, CalDAV, CardDAV and the console together, since they share a port. |
vsx_store_transaction_seconds | histogram | kind: transact, snapshot; outcome: ok, conflict, error | How long the database’s transactions take, as the server waited for them, retries included. |
vsx_queue_depth | gauge | queue: transport, delivery | Entries waiting, sampled every 15 seconds. |
vsx_queue_oldest_seconds | gauge | queue | Age in seconds of the entry that has been due longest, sampled every 15 seconds. |
A metric appears once something has been counted in it; a new node may show only some of them.
# HELP vsx_queue_depth Entries waiting in a queue.
# TYPE vsx_queue_depth gauge
vsx_queue_depth{queue="transport"} 3
vsx_queue_depth{queue="delivery"} 0
What to alert on
| Alert | Suggested rule | Why |
|---|---|---|
| Mail is stuck | vsx_queue_oldest_seconds{queue="transport"} > 3600 for 15 minutes | A deep queue that moves is healthy; one whose oldest entry keeps ageing is not. |
| Delivery into mailboxes is stuck | vsx_queue_oldest_seconds{queue="delivery"} > 300 | No node is running the deliver role, or delivery is failing. |
| Node not ready | /readyz not 200 | The store is unreachable. |
| Password guessing | A sharp rise in rate(vsx_authentications_total{outcome="failed"}[5m]) | Someone is guessing; lockout is working, but you want to know. |
| Someone is harvesting addresses | Any vsx_sessions_closed_total{reason="network-refused"} | A network kept guessing at your addresses and is refused for an hour. The audit log’s receiving lines name it. |
| A filter stage is slow | A steady rise in rate(vsx_filter_stages_cut_total[15m]) | A blocklist or a resolver answers slowly, and scoring is being finished at delivery. |
| Mail to a provider is failing | A rise in rate(vsx_outbound_recipients_total{outcome="deferred",provider="google.com"}[15m]) against its delivered rate | A provider is putting your mail off. Sending and delivery explains its pace. |
| Mail leaving unencrypted | Any rise in vsx_outbound_connections_total{tls="none"} | Some receiving servers offer no encryption. A rise for a named provider is worth a look. |
| Directory unreachable | Any vsx_authentications_total{outcome="unavailable"} | People who sign in with their directory password cannot until it answers. Check the directory and the network to it. |
| Something needs fixing | versealx-server doctor exits with 2 | See below. Run it every hour or so. |
Alerts to an organisation’s administrators
An organisation’s administrators can be told by mail when something needs them. Each rule names what to watch, the number at which to go off, and whom to tell:
| Watching | Counts | Goes off at |
|---|---|---|
queue-backlog | The organisation’s messages waiting to be delivered | A number from 1 to 100,000 |
dmarc-failures | Messages in the organisation’s name that failed DMARC, in the reports received for the last day | A number from 1 to 1,000,000,000 |
sign-in-failures | Accounts locked out, or collecting failed sign-ins now | A number from 1 to 100,000 |
dns-problem | Domains whose last DNS check found something broken | 1 |
breached-passwords | Accounts whose password in use has appeared in a data breach and must be changed; see Passwords that have leaked | A number from 1 to 100,000 |
passkey-went-back | Passkeys whose signature counter went backwards in the last day, which is what a copied authenticator does; see Passkeys | 1 |
forwarding-requests | Requests to forward mail outside the organisation that wait for an administrator; see Forwarding outside the organisation | A number from 1 to 100,000 |
sending-held | An account’s sending put on hold for review because it sent unlike itself, each time it happens; see Sending held for review | 1 |
organisation-storage | The organisation’s mailboxes at or over its storage warning level, before its ceiling refuses mail for everybody; see The organisation’s storage | 1 |
reported-phishing | A message somebody reported as phishing, or a later verdict found, and each one the server took back from everybody, each time it happens; see Reported phishing | 1 |
certificate-expiry | A certificate for the organisation’s domains running out within 14 days without having been renewed, each certificate on its own, and again at 3 days, with why its last renewal failed | 1 |
blocklisted | One of the organisation’s domains on a blocklist the server looks at, each listing on its own, and again when it clears; see Blocklists | 1 |
mail-stuck | The organisation’s messages waiting more than four hours for a server that answers but will not take them, which usually means a policy block there | A number from 1 to 100,000 |
bounces-rising | The percent of the organisation’s messages sent out today that bounced for at least one recipient, once it has sent 20 | A percent from 1 to 100 |
approvals | Each request that starts waiting for a second administrator’s approval, and each time the server’s operator goes ahead without asking, as an emergency | 1 |
account-recovered | Each account recovered with its recovery address, when somebody forgot their password; see Forgotten passwords | 1 |
webhook-paused | Each webhook, or the security log export, paused after failing for a day | 1 |
break-glass-used | Each sign-in of a break-glass account | 1 |
below-strict | The organisation’s security controls that fall short of the Strict profile. Not among the rules an organisation starts with | A number from 1 to 30 |
new-networks | The accounts that signed in from a network they hadn’t used in ninety days, in the last hour; see A sign-in from somewhere new | A number from 1 to 100,000 |
not-me | Each time somebody says a sign-in to their account was not them | 1 |
mail-damaged | Messages the integrity scrub found lost or damaged: those it couldn’t repair, and those it restored from a backup in the last day | A number from 1 to 100,000 |
journal-delayed | The organisation’s journal reports that have not reached their archive after a day | A number from 1 to 100,000 |
lookalike-domain | A newly found look-alike of the organisation’s domains, once per domain | 1 |
An organisation that has never saved its rules has thirteen, sent to its administrators: passkey-went-back, sending-held, organisation-storage, reported-phishing, approvals, account-recovered, webhook-paused, break-glass-used, not-me, journal-delayed, mail-damaged and lookalike-domain, each at 1, and new-networks at 20. An alert goes off when its number reaches the threshold, and not again until the number has fallen below it and reached it again. Each time, everybody the rule names is sent one plain-text message saying what was watched, the number now, the threshold, and where to look. Every alert is also kept in a history for 90 days.
Write the rules as a file and set them:
{
"rules": [
{ "when": "queue-backlog", "threshold": 500, "notify": ["postmaster@example.com"] },
{ "when": "dns-problem", "threshold": 1, "notify": ["it@example.com", "on-call@pager.example.net"] }
]
}
vsx admin alerts set alerts.json
vsx admin alerts show
vsx admin alerts history
The server’s own alerts
The server has alerts of its own, for what only it can see. They are tenant 0’s, and only its operator sets them:
| Watching | Counts | Goes off at |
|---|---|---|
certificate-expiry | Every certificate the server serves, its own and each domain’s | 1 |
blocklisted | The server’s sending addresses and every organisation’s domains, each listing on its own | 1 |
mail-stuck | Every organisation’s stuck messages | A number from 1 to 100,000 |
disk-filling | Each volume a node keeps its store or its mail on, at this percent full or more | A percent from 50 to 99 |
bounces-rising | The percent of every organisation’s messages sent out today that bounced | A percent from 1 to 100 |
provider-throttling | A receiving provider (Google, Microsoft, Yahoo, iCloud) asking the server to slow down for an hour or more | 1 |
backup-failed | The node’s daily backup failing, or none read back whole in 26 hours | 1 |
slow-scanner | Runs of a scanner or milter that took more than half the filter’s time budget, in the last hour | A number from 1 to 100,000 |
backup-test-failed | The weekly restore test of the newest backup failing | 1 |
node-missing | A node of the cluster not heard from for 30 seconds, each time it goes missing; see Nodes of a cluster | 1 |
probation-kept | A new organisation’s probation kept past its days because too much of its outside mail bounced, once for each organisation | 1 |
disk-filling, provider-throttling, backup-failed, backup-test-failed, node-missing, probation-kept and slow-scanner are the server’s alone: an organisation that asks for one is refused. Until the operator saves otherwise, the server has all eleven: disk-filling at 85, bounces-rising at 5, slow-scanner at 10, and the rest at 1. A node that has never been set up to take daily backups is not told about them. They tell nobody by mail until the operator names whom, and each is logged as a warning, and kept in the history, either way. Set them over the local socket, as tenant 0:
vsx admin alerts show --tenant 0
vsx admin alerts set server-alerts.json --tenant 0
The server’s alerts go every way an organisation’s go. As tenant 0, the operator can add a webhook that takes alert events, or export the security log to a SIEM. Each server alert is then sent there as alert.raised, even when it tells nobody by mail. On the console, the operator does this from the Alerts and Webhooks pages. From the command line:
vsx admin webhooks add https://pager.example.com/hooks/mail --events alert --description 'On call' --tenant 0
Blocklists
Once an hour, one node looks the server up on blocklists: each of its sending addresses (what its host name resolves to, leaving out private ones) on the lists of addresses, and each organisation’s domains, up to 1,000, on the lists of domains. The lists are the filter’s own blocklists, leaving out its allowlists. Where the filter names no list of a kind, Spamhaus ZEN is used for addresses and Spamhaus DBL for domains.
The lookups go through the node’s own resolver, as the lists’ terms require. A list asked through a public resolver refuses to answer; a refusal is never counted as a listing, and doctor says which list refused and why.
A rule may name up to 10 addresses, and an organisation may have up to 20 rules. A rule with no addresses is kept in the history only. Administrators set the rules; auditors can see them and the history. Over the API they are GET and PUT /api/v1/tenants/{tenant}/alerts, the PUT with the version the GET gave, and GET /api/v1/tenants/{tenant}/alerts/history.
What is stored
Once a day, one node counts what each organisation stores, by kind:
- message content (each message counted once, however many people hold it)
- message records
- calendars and contacts
- quarantine
- message traces
- the audit log
- reports
- change logs
The count is paced so that it never competes with mail. Each node also records how full its disks are once a day. From 30 days of those records, the server forecasts the day each disk will be full, when it is growing.
versealx-server storage status
versealx-server storage organisations
versealx-server storage kinds --tenant 3
The operator reads the same on the console’s Storage page, beside Usage, and through GET /storage, GET /storage/organisations and GET /storage/kinds. An organisation’s administrators and auditors see their own counts on the What is stored card on the Usage page, with versealx-server admin usage breakdown, or through GET /tenants/{tenant}/storage/breakdown. Counts are kept for 400 days.
When a disk fills
A node never damages mail or bounces it because a disk filled. When a volume holding the store or the blobs reaches the reserve ([storage] reserve_percent, 95% unless you change it), the node stops taking new mail. MAIL FROM on port 25, on submission and from your other premises gets this reply:
452 4.3.1 Insufficient system storage; try again later
That is a temporary refusal. Senders keep the message and try again, and mail apps keep it in their outbox. Everything else goes on. People can read, move and delete mail, and administrators can make room. As soon as the volume is back under the reserve, mail is accepted again without any action from you. The node checks its disks at most every ten seconds.
Well before the reserve, the server’s disk-filling alert tells you a volume is filling (85% unless you change it).
Logs
The server writes its log to standard output. The packaged service sends it to the systemd journal:
journalctl -u versealx-server -f
- Lines are human-readable, with a time, a level and named fields.
- The level is
infounless you set theRUST_LOGenvironment variable, for exampleRUST_LOG=debugin a systemd drop-in. - Every line about a message carries its queue id, and no line ever carries what a message says. Relay passwords are never written.
Lines worth knowing
| Line | Meaning |
|---|---|
mx listening, submission listening, imap listening, pop3 listening, managesieve listening | A listener is bound, with its address. |
https listener | The shared HTTPS listener is bound, with which services it carries. |
local socket listening | The admin socket is ready. |
metrics and health listening | The metrics endpoint is bound. |
relay worker started, deliver worker started | The queue workers are running. |
generated a new key-encryption key; back it up | First start: copy the key file now. See Backup and restore. |
accepted | A message was accepted, with its size, recipients and role. |
filtered | The filter’s verdict, the deciding stage and the score. |
authenticated, authentication failed | A submission sign-in, with the mechanism. |
no DKIM key; sending unsigned | A domain has no signing keys. |
certificate obtained, could not obtain a certificate; will try again, serving the certificate another node obtained | ACME. |
domain proved, domain no longer proves itself; the record it was proved by has gone | Domain ownership changed. |
yesterday's reports sent, could not send yesterday's reports | The daily DMARC and TLS reports. |
old message traces forgotten | Trace retention ran. |
sieve script failed; keeping | A person’s script failed; the message went to their inbox. |
address is not public; skipped | A destination resolved to a private address. |
the serve role is on but this node runs neither the store nor the submission role; … | Clients would be told to connect to services this node does not run. |
shutting down | The node received a stop signal. |
explain
versealx-server explain prints what a node’s configuration will do without starting it: each role, what it listens on and what it sends out of the machine; the outbound pacing; the message ceilings; where data lives; the filter pipeline; and what leaves the premises — certificate requests, metrics and whether the node may deliver to private addresses. See the Command-line reference.
doctor
doctor checks a deployment from the outside in and tells you what to do about each problem. It reads DNS, the certificate, the store and the clock, and changes nothing.
sudo -u versealx versealx-server doctor --config /etc/versealx-server/versealx-server.toml
ok hostname: mail.example.com
ok tls certificate: /var/lib/versealx-server/tls/certificate.pem
ok tls key: /var/lib/versealx-server/tls/key.pem
ok clock: reads 2026
ok resolver: passes DNSSEC through; a name that does not exist can be proved not to and the answer kept
ok reverse DNS: mail.example.com and its 1 address(es) confirm each other
ok DANE: no _25._tcp.mail.example.com record, so nothing is promised about this host's key
ok store: answers; 1 domain(s) configured
ok example.com MX: names this host (mail.example.com)
BAD example.com SPF: does not name the relay this node sends through: v=spf1 mx -all
mail leaves from relay.example.net, not from this domain's MX, so every message fails SPF and then DMARC. Publish `example.com. IN TXT "v=spf1 mx include:relay.example.net -all"`
ok example.com ownership: _versealx-verify.example.com is published, so autoconfig and MTA-STS are answered
ok example.com DMARC: v=DMARC1; p=none; rua=mailto:dmarc@example.com
…
16 checked, 15 good, 0 to look at, 1 broken
Each line starts with ok, warn (works, but will not keep working, or is not what you meant) or BAD (mail will be lost or refused). A problem’s remedy is on the indented line below it.
Exit codes
| Code | Meaning |
|---|---|
0 | Everything is ok. |
1 | Something is warn. |
2 | Something is BAD, or doctor could not run — for example the resolver could not start: <reason>, or the configuration could not be read. |
The checks
| Check | What it looks at |
|---|---|
hostname | A fully qualified name. warn for a single-label name or one ending in .localhost. |
tls certificate, tls key | With certificate files: both exist. BAD when one is missing. |
tls | With ACME: a certificate has been obtained and has more than 30 days left. See TLS certificates. |
tls renewal | For each certificate whose last order failed: when, and what the CA or the network said. warn. |
clock | The clock reads 2024 or later. |
resolver | Whether your resolver passes DNSSEC records through, by asking for a name that cannot exist. warn when it does not, times out, or invents an answer. |
reverse DNS | The host name resolves; each of its first two addresses has a PTR; the PTR resolves back. See DNS records. |
DANE | Any TLSA record for the host matches the certificate in use, and the state of a key rollover. See DNS records. |
store | The store opens and answers, and how many domains it holds. BAD when it does not; the DNS checks that need no domain still run. |
domains | warn when there are none. |
leaked-password check | Whether the leaked-password service answers, and how quickly, or that this node does not ask (breach_check = false). |
confinement | Whether the kernel enforces the node’s confinement: Landlock, version <n>, and anything this kernel cannot confine. BAD when landlock = "required" and the kernel does not enforce it, so run would refuse to start. |
key-encryption key | When the key was last rotated, or that it never has been. BAD when a rotation is half done: no node starts until the same keys rotate command is run again. |
<domain> MX | The MX names this host. |
<domain> SPF | A record exists and agrees with whether this node sends through a relay. |
<domain> ownership | The _versealx-verify token is published. |
<domain> DMARC | A DMARC record exists. |
<domain> DKIM <selector> | Each signing key’s record is published and matches the key in the store. |
<domain> MTA-STS | The advertised policy and the policy this node serves agree, and every MX host is covered. |
<domain> TLS reporting | A TLS reporting record exists. |
blocklist | The server’s sending addresses and its domains, looked up on the blocklists as the hourly look does. BAD for each listing, with the list’s page for asking to be removed; warn for a list that did not answer. |
A lookup that fails rather than finding nothing is reported as warn with check the resolver before reading anything else here.
Something unclear or out of date on this page? Tell us.