aslain.dev
0%
01 Hizmetler 02 Hakkımda 03 Projeler 04 Stack 05 Blog 06 İletişim
← Tüm makaleler Discord Bots

Discord Bot Uptime: Crash Detection and Auto-Restart

Even a well-written bot can quietly die at 3 a.m., which is why discord bot uptime is less about code and more about operational discipline. The problem is rarely that the bot is "bad" — it's that nobody notices when it crashes, and there's no mechanism to bring it back up. In this guide we'll set up, step by step, how to detect why a bot goes down, how to make it recover itself with pm2, and how to monitor it from the outside with a simple healthcheck.

Why do bots crash?

Know your enemy first. The overwhelming majority of Discord bots stop for a handful of recurring reasons:

  • Uncaught errors: An exception thrown inside an event handler that isn't wrapped in try/catch takes down the whole Node.js process.
  • Swallowed promise rejections: A promise that isn't awaited or given a .catch() will crash the process on modern Node versions.
  • WebSocket drops: Network blips or restarts on Discord's side break the connection; discord.js usually reconnects on its own, but can sometimes get stuck in a disconnect state.
  • Memory leaks: An ever-growing cache or uncleared setIntervals will get the process killed with OOM (Out Of Memory) within hours.
  • Rate limits and invalid tokens: If the token is reset or bad code triggers a flood of 429s, the gateway connection is refused.

Looking at this list, the solution becomes clear: make errors visible, put a watchdog in front of the process, and measure from the outside whether the bot is genuinely "alive."

Step 1: Stop swallowing errors

The first line of defense is the code itself. Even when the bot crashes, it has to log why it crashed, otherwise you'll loop the same error forever under restarts. Add listeners to the discord.js client and to the global process events:

client.on('error', (err) => console.error('[client error]', err));
client.on('shardError', (err) => console.error('[shard error]', err));

process.on('unhandledRejection', (reason) => {
  console.error('[unhandledRejection]', reason);
});

process.on('uncaughtException', (err) => {
  console.error('[uncaughtException]', err);
  process.exit(1); // exit deliberately; let pm2 restart
});

There's an important design decision here: when uncaughtException fires, we end the process on purpose. Trying to keep the process alive after an unknown error means continuing in a half-broken state. The clean approach is: log and exit, and leave the recovery to a process manager.

Step 2: Automatic restarts with pm2

pm2 is a process manager that runs Node apps in the background, restarts them automatically when they crash, and collects their logs. Install it globally on your VPS:

npm install -g pm2

You can start the bot directly, but keeping the configuration in an ecosystem.config.js file is much cleaner:

module.exports = {
  apps: [{
    name: 'discord-bot',
    script: './index.js',
    instances: 1,
    autorestart: true,
    max_memory_restart: '300M',
    restart_delay: 5000,
    max_restarts: 10,
    min_uptime: '10s',
    env: { NODE_ENV: 'production' }
  }]
};

The fields that matter:

  • autorestart: When the process exits, pm2 restarts it.
  • max_memory_restart: Restarts the bot proactively if memory exceeds 300 MB — insurance against leaks.
  • restart_delay: Waits 5 seconds before restarting, so a permanent error such as an invalid token doesn't create an endless loop that pins the CPU.
  • min_uptime + max_restarts: If the bot crashes 10 times without staying up for 10 seconds, pm2 marks it as "errored" and stops; this prevents you from hiding a real problem under endless restarts.

To start it and have it come back up after a server reboot:

pm2 start ecosystem.config.js
pm2 save
pm2 startup   # run the command it prints

Use pm2 logs discord-bot and pm2 status to watch logs and state.

Step 3: A healthcheck endpoint

pm2 answers the question "is the process running?", but it can't catch the "process is up but the Discord gateway has dropped" state. To verify the bot is truly alive, add a small HTTP healthcheck. The built-in http module is enough:

const http = require('http');

http.createServer((req, res) => {
  // client.ws.status === 0 -> READY
  const healthy = client.isReady() && client.ws.status === 0;
  res.writeHead(healthy ? 200 : 503, { 'Content-Type': 'application/json' });
  res.end(JSON.stringify({
    status: healthy ? 'ok' : 'degraded',
    ping: client.ws.ping,
    uptime: process.uptime()
  }));
}).listen(3001);

Now curl http://localhost:3001 shows the bot's gateway status and ping. It's critical to tie this endpoint to whether the gateway is actually READY, not just to process.uptime() — otherwise you'll mistake a "living dead" bot for a healthy one.

Step 4: External monitoring and alerts

The final link is being notified when something goes wrong. Hook the healthcheck endpoint up to a free monitoring service (UptimeRobot, Better Stack, or your own cron) so that when it returns 503 or stops responding, you get an email/Discord-webhook alert. A simple self-monitoring approach also works:

setInterval(() => {
  if (!client.isReady()) {
    console.error('[health] bot not READY, exiting');
    process.exit(1); // let pm2 do a clean start
  }
}, 60_000);

If the gateway gets stuck silently, this loop intentionally drops the process so pm2 can do a clean reconnect. It's usually better to wait for discord.js's own reconnection logic, but for "zombie" states this last resort is safe.

Frequently Asked Questions

Can I keep my bot running 24/7 on free hosting (like Replit)?

Most free tiers put the process to sleep when there's no traffic, so the most reliable path to true 24/7 uptime is a small VPS. Keep-alive ping tricks are fragile and violate many platforms' terms of service. For a single bot, a few-dollars-a-month VPS plus pm2 is far more dependable.

Should I use systemd or Docker instead of pm2?

All three are valid. systemd is built into Linux and offers service-level restarts; Docker does the same with a restart: unless-stopped policy. What makes pm2 attractive is that log collection, memory-based restarts, and ergonomics like pm2 status come in a single tool. If you already use Docker, the container restart policy is usually enough.

How do I stop a restart loop that keeps crashing?

min_uptime and max_restarts exist exactly for this: if the bot crashes too many times in a short window, pm2 gives up. At that point use pm2 logs to find the root cause (usually an invalid token, a missing env variable, or a hard error thrown in the code) and fix it.

Is your bot still dropping for no clear reason? I can set up crash detection, pm2 configuration, and healthcheck monitoring for you; for a resilient, self-healing bot infrastructure, get in touch with me.

Bu kategorideki tüm yazılar →

Devamı için