Networking troubleshooting
Networking troubleshooting is a core skill for DevOps engineers, because most production issues involve connectivity problems between services — such as:
Application cannot reach database
API is not responding
DNS is misconfigured
Port is not open
Network routing problems
Linux provides powerful tools to diagnose these issues quickly. Below is a deep dive into the most important networking commands used in DevOps and production debugging.
1️⃣ ping — Test Network Connectivity
What ping does
ping checks whether a remote host is reachable over the network.
It works using the ICMP protocol.
Command:
ping google.com
Example output:
64 bytes from 142.250.183.206: icmp_seq=1 ttl=117 time=24.3 ms
64 bytes from 142.250.183.206: icmp_seq=2 ttl=117 time=25.1 ms
Important Metrics
| Field | Meaning |
|---|---|
icmp_seq |
packet number |
ttl |
time to live |
time |
response time (latency) |
Stop ping:
Ctrl + C
DevOps Use Case
Check if server is reachable
Example:
ping 10.0.0.5
If unreachable:
Destination Host Unreachable
Possible causes:
server down
firewall blocking ICMP
network issue
2️⃣ curl — Send HTTP Requests
What curl does
curl is used to send HTTP requests and test APIs.
Command:
curl http://localhost:8080
Example response:
{"status":"running"}
DevOps Use Cases
Test API endpoints
curl https://api.example.com/users
Check HTTP status
curl -I https://example.com
Output:
HTTP/1.1 200 OK
Debug microservices
Example architecture:
Frontend → API → Database
Test API service:
curl http://api-service:8080/health
If response fails:
Connection refused
Service may not be running.
3️⃣ wget — Download Files
What wget does
wget downloads files from the internet via HTTP, HTTPS, or FTP.
Example:
wget https://example.com/file.zip
Output:
Saving to: ‘file.zip’
DevOps Use Cases
Download software packages
Example:
wget https://dl.k8s.io/release/v1.29/bin/linux/amd64/kubectl
Download logs or backups
Example:
wget https://server/backups/db_backup.tar.gz
Test file accessibility
If download fails:
404 Not Found
File may not exist.
4️⃣ netstat — Check Network Connections
What netstat does
netstat shows:
open ports
active connections
listening services
Command:
netstat -tulnp
Option Explanation
| Option | Meaning |
|---|---|
t |
TCP connections |
u |
UDP connections |
l |
listening ports |
n |
numeric addresses |
p |
process name |
Example output:
tcp 0 0 0.0.0.0:80 LISTEN 1023/nginx
tcp 0 0 0.0.0.0:22 LISTEN 500/sshd
Meaning:
| Port | Service |
|---|---|
| 80 | Nginx |
| 22 | SSH |
DevOps Use Case
Example issue:
Website not accessible
Check if port 80 is open:
netstat -tulnp | grep 80
If nothing appears → service not running.
5️⃣ ss — Modern Replacement for netstat
ss is faster and more powerful than netstat.
Command:
ss -tulnp
Example output:
LISTEN 0 128 0.0.0.0:80 users:(("nginx",pid=1023))
DevOps Use Case
Check if application is listening on port:
ss -tulnp | grep 8080
Output:
LISTEN 0 128 0.0.0.0:8080 java
Means Java service running on port 8080.
6️⃣ traceroute — Trace Network Path
What it does
traceroute shows the path packets take through routers to reach a destination.
Command:
traceroute google.com
Example output:
1 192.168.1.1
2 10.10.1.1
3 172.217.169.14
Each line is a router hop.
DevOps Use Case
Example problem:
Server reachable but slow
Run:
traceroute api.example.com
If a hop has high latency:
200 ms
Network routing problem exists.
7️⃣ dig — DNS Lookup Tool
dig queries DNS servers.
Command:
dig google.com
Example output:
google.com. 300 IN A 142.250.183.206
Meaning:
| Field | Meaning |
|---|---|
| google.com | domain |
| A | record type |
| IP | resolved address |
DevOps Use Cases
Verify DNS resolution
dig api.company.com
Query specific DNS server
dig @8.8.8.8 google.com
This checks Google DNS directly.
8️⃣ nslookup — DNS Query Tool
nslookup is another tool to query DNS.
Command:
nslookup google.com
Output:
Name: google.com
Address: 142.250.183.206
DevOps Use Case
Check DNS misconfiguration.
Example problem:
API domain not resolving
Run:
nslookup api.example.com
If result:
server can't find api.example.com
DNS record missing.
🚀 Real DevOps Debugging Scenario
Problem
Users cannot access website.
Step 1 — Check connectivity
ping example.com
Step 2 — Verify DNS
dig example.com
Step 3 — Check service
ss -tulnp | grep 80
Step 4 — Test HTTP response
curl http://example.com
Step 5 — Check network route
traceroute example.com
This workflow helps isolate the problem quickly.
📊 Summary
| Command | Purpose |
|---|---|
ping |
test connectivity |
curl |
test HTTP/API |
wget |
download files |
netstat |
view network connections |
ss |
modern netstat |
traceroute |
trace network path |
dig |
DNS lookup |
nslookup |
DNS query |
⭐ Why These Tools Matter in DevOps
DevOps engineers use these commands to:
debug microservice communication
verify API availability
diagnose DNS issues
check open ports
analyze network routes
validate infrastructure connectivity
These tools are essential for troubleshooting distributed systems and cloud infrastructure.
A real production incident often involves multiple services failing to communicate with each other. In a microservices architecture, one service depends on several others (API, database, authentication service, etc.). If connectivity breaks anywhere, the entire application may fail.
Below is a step-by-step example of how DevOps engineers debug a microservice connectivity issue in production.
🚨 Real Production Incident: Debugging a Microservice Connectivity Issue
Scenario
A company runs a web application with this architecture:
User → Nginx → API Service → Database
Suddenly users start reporting:
500 Internal Server Error
The website is not working.
The DevOps engineer needs to identify where the failure occurs.
Step 1️⃣ Verify the Website Response
First, check if the service responds.
Command:
curl http://example.com
Response:
HTTP/1.1 500 Internal Server Error
This confirms the server is reachable but something inside the system is failing.
Step 2️⃣ Check Nginx Logs
Logs usually show the first hint.
tail -f /var/log/nginx/error.log
Example error:
connect() failed (111: Connection refused) while connecting to upstream
This means:
Nginx → cannot connect to backend API
So the problem is likely with the API service.
Step 3️⃣ Check If the API Service Is Running
Check if the service port is listening.
ss -tulnp | grep 8080
Example output:
LISTEN 0 128 0.0.0.0:8080 users:(("java",pid=2431))
If no output appears:
API service is not running
Restart service:
systemctl restart api-service
If service is running, move to next step.
Step 4️⃣ Test API Directly
Now test the backend API.
curl http://localhost:8080/health
Example output:
{"status":"DOWN"}
This indicates the API itself has an internal problem.
Step 5️⃣ Check API Logs
Look at application logs.
tail -f /var/log/app.log
Example output:
Database connection failed
Now we know the issue is between:
API → Database
Step 6️⃣ Verify Database Connectivity
Test database port.
Example MySQL:
ss -tulnp | grep 3306
If nothing appears:
Database service stopped
Start database:
systemctl start mysql
Step 7️⃣ Check DNS Resolution
If the API connects using a hostname:
db.internal.company
Check DNS:
dig db.internal.company
Example failure:
NXDOMAIN
DNS record missing.
Step 8️⃣ Check Network Connectivity
Test connectivity between services.
ping db.internal.company
Or test port access:
nc -zv db.internal.company 3306
Example result:
Connection refused
Database may be blocked by firewall.
Step 9️⃣ Check Firewall Rules
Example:
sudo ufw status
Or cloud firewall rules (AWS security groups).
If port is blocked:
Allow port 3306
Step 🔟 Confirm System Is Fixed
Restart services:
systemctl restart nginx
systemctl restart api-service
Test again:
curl http://example.com
Output:
HTTP/1.1 200 OK
System working again.
🔍 Root Cause Analysis
Final root cause:
Database service crashed → API could not connect → Nginx returned 500 error
📊 Typical DevOps Debugging Flow
DevOps engineers usually follow this troubleshooting chain:
Client → Website response
Web server logs
Backend service health
Application logs
Database connectivity
DNS resolution
Network routes
Firewall rules
Each step narrows down the exact failure point.
🚀 Tools Used in This Investigation
| Tool | Purpose |
|---|---|
curl |
test HTTP service |
tail |
monitor logs |
ss |
check open ports |
dig |
DNS lookup |
ping |
test connectivity |
systemctl |
manage services |
⭐ Key DevOps Lesson
In distributed systems:
Most outages are not code bugs.
They are connectivity problems between services.
Common causes include:
crashed service
wrong port configuration
DNS misconfiguration
firewall blocking traffic
database failure
Understanding Linux networking and log analysis tool