A student pinged me last week. The site was down. He had already restarted nginx twice. It was still down, and he could not tell me whether the process had been running, which port it was bound to, or what the log said the first time it died.
That is the whole problem. A restart feels like progress. Most of the time it throws away the only evidence you had.
Why restarting first makes the next hour longer
When a Linux service fails, the useful facts are short-lived. systemctl status still shows the last exit code. The journal still has the line from the crash. The port is still held by whatever grabbed it. The moment you restart, three things happen.
The process id changes, so the old one is gone. If the new start succeeds, status looks healthy and the failure is no longer the first thing you see. If it fails again, you now have two failures mixed together, and you cannot tell which lines belong to the first one.
I am not against restarting. I am against restarting before you have read the box. In the lab, and on a real server, the order is the same.
Check whether the process is running
Start with the service you think is broken. On a systemd machine that is usually:
systemctl status nginx --no-pagerRead the first lines, not the whole scrollback. You want four things: Active, the main PID, the last exit code, and the time it entered that state.
You get two outcomes, and they ask for different next moves.
The service is dead. Active says failed or inactive. A restart might bring it back, but you still do not know why it died. If you start it now and it dies again, you have learned nothing, and you have pushed the original lines further up the journal. Leave it stopped until you have read the log.
The service is running. Active says running, and there is a real PID. A restart will not fix a full disk, a bad upstream, or a process that is up and refusing work. Leave it alone. The fault is somewhere else, and restarting only adds noise.
If you are not on systemd, do not guess from a grep of the process list. This line lies:
ps aux | grep nginxThe grep matches itself, so you always see a line, even when nginx is not there. Ask for the master process by name:
ps -C nginx -o pid,user,cmdEmpty output means it is not running. A line with nginx: master process means it is. Worker processes under it are normal. One extra line that is only your grep is not a service.
Check the port the browser is actually using
A process can be running and still not be the thing answering the page. This is the check I lose the most lab time to.
ss -lntp-l is listening sockets, -n keeps port numbers numeric, -t is TCP, and -p prints the process name. You need root, or sudo, to see process names you do not own.
Look for the port in the browser. For a normal site that is 80 or 443. For a lab app it might be 8080 or 3000. The sheet says which. Do not assume.
Match three things. The port in the address bar. The port in the ss output. The process name beside it. If they do not match, you are editing a config the running process is not reading.
The classic version in class looks like this. Apache is bound to port 80. Nginx failed to start because that port was taken. The browser is talking to Apache. Every change you make in /etc/nginx is invisible, and another restart of nginx fails the same way.
LISTEN 0 511 0.0.0.0:80 0.0.0.0:* users:(("apache2",pid=912,fd=4))There is no nginx on 80. Restarting nginx again will not move Apache off that port.
Read the log from the failure, not from the end
Now, and only now, read what it said when it died.
journalctl -u nginx -n 80 --no-pagerEighty lines, not the whole journal. -u nginx keeps you in that unit. If you drop the -u, you will be reading ssh, cron, and the kernel, and you will miss the one line that matters.
Find the timestamp of the failure and read up from there. The newest lines are often your own restarts. The useful line is older, and it is boring and specific. These are the ones I actually see:
Address already in use. Something else owns the port. Go back to
ss -lntpand name that process. Do not start nginx again until the port is free, or until you have pointed nginx at a different port on purpose.Permission denied on a certificate or a log directory. The process user cannot read the file. Fix the path or the mode. A restart repeats the same denial.
No space left on device. Status may still say the unit is up, or it may have failed while rotating a log. Check
df -hbefore you touch the service. Restarting a full disk does not make space.A syntax error, with a file path and a line number. Open that file at that line.
nginx -twill tell you the same thing without starting anything. Fix the line, test again, then start once.
If the log is quiet and the process is up, the fault is probably not inside that unit. Think upstream. A proxy can be healthy while the app behind it is not, and the browser will still say the site is down.
One job from the lab, in order
The site on the lab VM would not load. The student had restarted nginx twice. Here is what the three checks actually returned.
Status said failed, with an exit code, not a running PID:
Active: failed (Result: exit-code)
Process: ExecStart=/usr/sbin/nginx (code=exited, status=1/FAILURE)The port check:
sudo ss -lntp | grep ':80'showed apache2, not nginx.
The log, above the two restart attempts, said:
nginx: [emerg] bind() to 0.0.0.0:80 failed (98: Address already in use)The fix was not a third restart. Apache was the program on port 80, left over from an earlier exercise. We stopped apache2, confirmed with ss -lntp that nothing was listening on 80, ran nginx -t, and started nginx once. Status went to running, and the page loaded.
If we had restarted nginx a third time without reading, we would have added another failure line and still served Apache to the browser.
The mistakes that waste the hour
Restarting in a loop. Each start without a reading buries the original error. If you have already restarted, say so, and look further back than the last minute.
Trusting grep. ps aux | grep nginx prints the grep. That is not nginx. Use ps -C or look for the master process.
Reading the wrong unit. The unit name is not always the package name you remember. On Debian the web server unit may be nginx or apache2. This is slower than guessing, and it is right:
systemctl list-units --type=service | grep -E 'nginx|apache'Changing three things, then starting once. If it comes up, you do not know which change did it. If it stays down, you do not know which change made it worse. One change, then look again.
What to do the next time a service is down
Run
systemctl statuson the unit and write down Active, the PID, and the exit code.Run
ss -lntpand confirm the port in the browser belongs to that process.Run
journalctl -ufor that unit and read the line from the failure, not the line from your restart.Change one thing. Test the config if the program has a test command. Start it once.
The same order works for any unit, not just nginx. Swap the name. The questions do not change.
If you want the wider set of checks I give the helpdesk class, they are in 15 Linux commands every helpdesk technician needs. We walk through this on a VM, service by service, in Introduction to System Administration.