Postmortem

Incident response: brute-force intrusion on the VPS (8-Sep-2026)

Blameless format. Times in UTC, taken from the proxy logs and the response session.

Summary and impact

On 7-Sep-2026 an attacker brute-forced the password of an administrator account on a WordPress staging site running on my VPS, uploaded their own plugins and dropped 21 executables. For about 15 hours they consumed CPU and RAM until the machine became unusable. The intrusion never left the site's container. Data was rescued, the server was rebuilt from scratch on a new machine, the configuration was hardened and the exposed credentials were rotated. Core services were back in under 5 hours from detection.

Impact

  • Everything on the VPS went down: the automation pipeline, n8n, the control panel and the demo sites.
  • Nothing running outside the VPS was affected.
  • No n8n data loss: its database was rescued intact.
  • Two personal deployments that existed only as builds copied to the server could not be rescued.

Duration

of undetected intrusion
15 h 42 min
from detection to containment
54 min
until n8n was back
4 h 52 min
until the last service
6 h 43 min

Timeline (UTC)

  1. Intrusion

    Seven failed attempts on wp-login.php, preceded by automated probing of other routes

  2. Intrusion

    Access: the password falls in 13 seconds

  3. Intrusion

    Straight to the plugin upload screen

  4. Intrusion

    Three foreign plugins and 21 hidden executables appear in core folders

  5. Detection

    Detection: the whole server is reported down

  6. Detection

    A single SSH login manages to measure load of 120 to 157 on 1 vCPU, 53 MB of free RAM and full swap

  7. Containment

    Containment: machine powered off, snapshot kept as evidence

  8. Containment

    Five-question forensic plan, all read-only

  9. Containment

    Forensics from the recovery environment, disk mounted and the compromised system never booted: vector, scope and timeline

  10. Recovery

    n8n database rescued (integrity_check = ok) along with config that wasn't in git

  11. Recovery

    n8n and the pipeline back on a new server

  12. Recovery

    From a temporary copy of the snapshot: clean-check of the other sites and copy of their volumes; the affected site is not restored

  13. Recovery

    The affected site rebuilt from scratch from its repository

  14. Hardening

    Hardening applied and verified

  15. Recovery

    Control panel back, now as a container

Detection and root cause

Detection

No alert caught it: the machine sat at 100 % for about 12 hours before anyone noticed. The alert that did exist reported the symptom (service down), not the cause.

Root cause

The staging site's wp-login.php was exposed to the internet with no rate limit on login attempts, and the site had an administrator account with a predictable name and a weak password. Seven attempts were enough. There was no vulnerability in our own code, WordPress core or the project's plugins.

Factors that widened the impact

  • All twelve sites shared a single database user with rights over all twelve databases.
  • No container had a memory limit: one site ate the whole machine's RAM and took n8n down with it.
  • The staging site was open to the internet and shared a machine with the pipeline.

Containment

  • Machine powered off and snapshot kept as evidence; the compromised system was never booted again.
  • None of the intruder's files were executed, not even to inspect them.
  • Scope verified, not assumed: cron, systemd, system users and authorized keys unchanged; no container mounted the Docker socket, so there was no path to the host; the other eleven sites were clean.

Recovery

  • Data rescue: the n8n database (15.7 MB plus its WAL), proxy and compose config, the sites' themes (251 MB, bind-mounted and outside the volumes) and the pipeline's workspace (77 MB).
  • Rebuild from scratch: a new server; n8n restored with its rescued database, its 2 workflows and 2 credentials; certificates issued automatically by Caddy.
  • The other ten sites were copied only after passing a clean-check: zero hidden files in core, zero PHP in uploads, zero foreign plugins.
  • The affected site was not restored: its database was dropped (with a prior copy stored off the server) and it was rebuilt from its repository.

Hardening

Applied on 8-Sep (verified on the running system, not on the file)

BeforeAfter
One database user with rights over all 12 databasesOne user per site, with rights only on its own; the shared one dropped
No memory limit384 MB per site and 512 MB for the database
Plugins could be uploaded and installed from the adminDISALLOW_FILE_MODS and DISALLOW_FILE_EDIT on; automatic minor core updates
Administrator account with a predictable nameThe rebuilt site has no such account; new 26-character password

Applied afterwards

  • Rate limit on wp-login.php
  • Load and disk alerts
  • After two memory-related outages (13 and 15 September): the panel requires at least 400 MB free and load below 2 before starting a site, with at most 2 running
  • The agent runner gained an inactivity watchdog

Pending

  • Second factor for WordPress administrators
  • Staging behind a password or IP allowlist

Credential rotation (driven by evidence, not reflex)

Rotated
The database credential the sites shared, assumed read because the attacker had arbitrary PHP in the container; the affected site's passwords.
Not rotated, with reasons
SSH keys (only public keys were on the server and they were unchanged); the Telegram bot token (the host's process list wasn't visible from inside the container); the database root password (it was never in the WordPress containers).

Retrospective

What went well

  • The compromised system was never booted: forensics ran with the disk mounted from the recovery environment.
  • Scope was proven with concrete checks before deciding what to rebuild.
  • The n8n database came out intact and the pipeline was back the same day.
  • Nothing was copied to the new server without passing the clean-check first.

What didn't go well

  • About 15 hours without detecting the intrusion: there were no load or disk alerts.
  • The problem was first attributed to 31-day-old zombie processes, which were a separate, older issue.
  • Rotating the Telegram token was requested before checking whether it was exposed; corrected after verifying.
  • The first rescue missed the sites' volumes and the bind-mounted themes; they were recovered before the temporary copy was destroyed.
  • Starting all ten sites at once pushed the new machine to 1,935 of 1,967 MB and load 21, and n8n restarted: the incident's failure mode, reproduced without an intruder.
  • The new server ended up without swap, which explained later outages. It has 2 GB of swap again.

Preventive actions

#ActionStatus
1One database user per site, with rights only on its ownDone
2Memory limit per containerDone
3DISALLOW_FILE_MODS on the sitesDone
4No administrator account with a predictable name; long, unique passwordsDoneOn the rebuilt site
5Start-up gated on memory and load; at most 2 sites at onceDone
6Rate limit on wp-login.phpDone
7Second factor for WordPress administratorsPending
8Staging behind a password or IP allowlistPending
9Load and disk alertsDone
10Separate the pipeline from exposed sitesAccepted riskMitigated with a per-service memory limit and at most 2 demos running
11Recreate the 2 GB swapDone

Lessons

  1. A staging site isn't a toy: it runs the same stack as production and deserves the same defenses, or shouldn't be on the open internet.
  2. If one credential spans twelve systems, one compromised site is worth twelve.
  3. Without a memory limit, any container can take the whole machine down.
  4. Alert on the cause (load, memory, disk), not just the symptom.
  5. Rotate what the evidence says was exposed, and document why the rest wasn't.
  6. Everything running on a server must be rebuildable from a repository.