पाठ 21 / 25

Logging, Metrics and Troubleshooting

Log useful fields, expose basic metrics and debug configuration problems.

Seeing what NGINX sees

The access log records each request; define a richer log_format with timing fields: $request_time (total time NGINX spent), $upstream_response_time (time the backend took), $upstream_addr (which backend answered), $upstream_cache_status and $request_id (a unique ID you can also forward to the app). JSON logs (log_format ... escape=json) are easy for log platforms to parse. The error log level (warn, error, info, debug) controls verbosity; errors such as "upstream timed out" or "connect() failed" pinpoint backend problems. The stub_status module exposes basic counters (active connections, accepted and handled connections, requests) for Prometheus exporters; restrict it to internal addresses. For troubleshooting: nginx -T shows the effective config; curl -v -H "Host: example.com" http://127.0.0.1/path tests a server block directly; compare $request_time and $upstream_response_time to tell whether slowness is in the app or between client and NGINX.

JSON access logs and a status endpoint

Timing fields show where time is spent; stub_status feeds metrics.

log_format json_combined escape=json
  '{"time":"$time_iso8601","request_id":"$request_id",'
  '"remote_addr":"$remote_addr","host":"$host","method":"$request_method",'
  '"uri":"$request_uri","status":$status,"bytes":$body_bytes_sent,'
  '"request_time":$request_time,"upstream_time":"$upstream_response_time",'
  '"upstream":"$upstream_addr","cache":"$upstream_cache_status",'
  '"user_agent":"$http_user_agent"}';

access_log /var/log/nginx/access.json json_combined;
error_log  /var/log/nginx/error.log warn;

server {
    listen 127.0.0.1:8081;
    location = /nginx_status {
        stub_status;
        allow 127.0.0.1;
        deny all;
    }
}

# proxy_set_header X-Request-ID $request_id;   # correlate with application logs

Upstream time tells you who is slow

If request_time is high but upstream_response_time is low, the delay is between NGINX and the client (slow network, large responses). If both are high, the application is slow.

त्वरित जाँच: Which variable shows how long the backend took to respond?

  • $request_time
  • $remote_addr
  • $upstream_response_time
  • $status
Answer

$upstream_response_time — $upstream_response_time measures the backend; $request_time covers the whole request.