The Slack Bot That Became Load-Bearing Infrastructure

A team built a quick Slack bot for deploy notifications. Eighteen months later, it ran deploys, managed on-call rotation, provisioned environments, and tracked incidents. It had no tests, no docs, and ran on a single EC2 instance nobody monitored.


I joined a client in August and within the first hour noticed something odd. Every conversation in their engineering Slack — about deploys, rollbacks, environment setup, on-call scheduling — ended with someone typing a command that started with /ops. Deploy to staging? /ops deploy staging. Need a review environment? /ops env create ticket-4821. Check who's on call? /ops oncall.

I asked the tech lead what /ops was. She said "the bot." I asked where the bot lived. She paused. "I think it's on Marcus's instance."

Marcus left four months ago.

How it started

The story, as I pieced it together from git blame and Slack history, went like this. About eighteen months earlier, a backend engineer named Marcus got tired of SSHing into a bastion host to trigger deploys. He wrote a small Node.js app — maybe 200 lines — that listened for Slack slash commands and ran deploy scripts via SSH. He put it on a t3.micro EC2 instance, added a systemd service file, and told the team about it in standup.

People loved it. No more bastion host. No more remembering which script to run for which environment. Just type /ops deploy staging and watch the output stream into a Slack thread.

That should have been the end of it. But convenience is a powerful force.

Within a month, someone asked Marcus to add rollback support. Then deploy notifications — a message in #releases whenever a deploy finished. Then someone on the platform team asked if the bot could provision short-lived review environments, because the existing Terraform workflow took 15 minutes of manual steps. Marcus added that too. Then on-call rotation management, because PagerDuty's Slack integration was clunky and the bot already had everyone's user IDs mapped.

By the time Marcus left in April, the bot was about 3,400 lines of JavaScript spread across 22 files. It had one dependency on a Slack SDK, one on the AWS SDK, and a local SQLite database for tracking environment state and on-call schedules. It had zero tests. The README said "run node index.js."

The day it went down

September 3rd, a Tuesday. A developer typed /ops deploy staging and got no response. She tried again. Nothing. Tried /ops ping — a health check someone had added. Silence.

The EC2 instance's root volume was full. CloudWatch wasn't set up, so there were no alarms. The bot had been writing deploy logs to a file that grew unbounded. Eight months of deploy logs, plus SQLite WAL files that were never cleaned up, ate the entire 8GB disk.

Here's the part that made it a real incident: nobody could deploy. The CI pipeline produced Docker images, but the actual deployment — the kubectl apply, the health check, the traffic shifting — was all wired through the bot. There was no other path to production.

The team spent 40 minutes SSHing into the instance, clearing logs, and restarting the process. During those 40 minutes, a hotfix for a payment calculation bug sat in a Docker registry, ready to go, while customers were being charged incorrect amounts.

What we actually found

When I dug into the bot's code, the technical problems were predictable: no log rotation, no health checks, no restart-on-failure, single point of failure, credentials stored in a .env file on the instance. Standard stuff.

The interesting problem was organizational. The bot had become what I'd call shadow infrastructure — critical systems that exist outside the team's mental model of their architecture. Ask anyone to draw the system architecture and they'd show you Kubernetes clusters, RDS instances, Redis, an API gateway. Nobody would draw the bot. It wasn't in any architecture diagram. It wasn't in any runbook. It wasn't in the infrastructure-as-code repository.

But it was the single most-used piece of internal tooling the team had. I checked the Slack logs: the team ran an average of 34 bot commands per day. Deploys, environment management, on-call queries, incident tracking. More interactions per day than their CI system.

Warning

If your team has built tooling that isn't in your infrastructure repo, isn't monitored, and isn't documented — but people use it every day — you have shadow infrastructure. The question isn't whether it'll cause an incident. It's when.

The fix (and the argument)

There was a genuine debate about what to do. One camp wanted to rewrite the bot properly — tests, CI/CD, run it in Kubernetes, add monitoring. The other camp wanted to kill it entirely and use off-the-shelf tools for each capability: GitHub Actions for deploys, PagerDuty's native integration for on-call, Terraform Cloud for environments.

We went with a middle path. We moved the deploy logic back into the CI pipeline where it belonged — GitHub Actions workflows triggered by merge or manual dispatch. We kept the Slack interface for visibility (deploy notifications, on-call lookups) but made it read-only. The bot could tell you things, but it couldn't do things. The doing happened in systems that were monitored, version-controlled, and understood.

The migration took about three weeks. The hardest part wasn't technical. It was convincing people to give up the convenience. /ops deploy staging is faster than navigating to GitHub, finding the workflow, clicking "Run workflow," selecting the environment, and clicking "Run." The developers were right about that. But speed of the happy path doesn't matter when the unhappy path is "nobody can deploy for 40 minutes."

The pattern

I've seen this pattern at four different clients now. The details vary — sometimes it's a Slack bot, sometimes it's a CLI tool someone wrote, sometimes it's a Jupyter notebook that generates reports everyone relies on. The shape is always the same:

  1. Someone builds a small tool to scratch an itch
  2. It works well, so people start using it
  3. Features accrete because requests are small and the builder is helpful
  4. The builder moves on (leaves, changes teams, gets busy)
  5. The tool is now critical but has no owner, no tests, and no operational story
  6. Something breaks and everyone discovers how dependent they were

The uncomfortable truth is that step 2 is where the intervention should happen. Not by killing the tool — that just punishes the person who built something useful. But by asking: if this works, where does it live? Who owns it? What happens when it breaks at 2 AM?

Most teams never ask those questions until the answer is already "we don't know."