Skip to main content

Command Palette

Search for a command to run...

Networking troubleshooting

Updated
•9 min read•View as Markdown

Networking troubleshooting is a core skill for DevOps engineers, because most production issues involve connectivity problems between services — such as:

  • Application cannot reach database

  • API is not responding

  • DNS is misconfigured

  • Port is not open

  • Network routing problems

Linux provides powerful tools to diagnose these issues quickly. Below is a deep dive into the most important networking commands used in DevOps and production debugging.

1️⃣ ping — Test Network Connectivity

What ping does

ping checks whether a remote host is reachable over the network.

It works using the ICMP protocol.

Command:

ping google.com

Example output:

64 bytes from 142.250.183.206: icmp_seq=1 ttl=117 time=24.3 ms
64 bytes from 142.250.183.206: icmp_seq=2 ttl=117 time=25.1 ms

Important Metrics

Field Meaning
icmp_seq packet number
ttl time to live
time response time (latency)

Stop ping:

Ctrl + C

DevOps Use Case

Check if server is reachable

Example:

ping 10.0.0.5

If unreachable:

Destination Host Unreachable

Possible causes:

  • server down

  • firewall blocking ICMP

  • network issue


2️⃣ curl — Send HTTP Requests

What curl does

curl is used to send HTTP requests and test APIs.

Command:

curl http://localhost:8080

Example response:

{"status":"running"}

DevOps Use Cases

Test API endpoints

curl https://api.example.com/users

Check HTTP status

curl -I https://example.com

Output:

HTTP/1.1 200 OK

Debug microservices

Example architecture:

Frontend → API → Database

Test API service:

curl http://api-service:8080/health

If response fails:

Connection refused

Service may not be running.


3️⃣ wget — Download Files

What wget does

wget downloads files from the internet via HTTP, HTTPS, or FTP.

Example:

wget https://example.com/file.zip

Output:

Saving to: ‘file.zip’

DevOps Use Cases

Download software packages

Example:

wget https://dl.k8s.io/release/v1.29/bin/linux/amd64/kubectl

Download logs or backups

Example:

wget https://server/backups/db_backup.tar.gz

Test file accessibility

If download fails:

404 Not Found

File may not exist.


4️⃣ netstat — Check Network Connections

What netstat does

netstat shows:

  • open ports

  • active connections

  • listening services

Command:

netstat -tulnp

Option Explanation

Option Meaning
t TCP connections
u UDP connections
l listening ports
n numeric addresses
p process name

Example output:

tcp 0 0 0.0.0.0:80 LISTEN 1023/nginx
tcp 0 0 0.0.0.0:22 LISTEN 500/sshd

Meaning:

Port Service
80 Nginx
22 SSH

DevOps Use Case

Example issue:

Website not accessible

Check if port 80 is open:

netstat -tulnp | grep 80

If nothing appears → service not running.


5️⃣ ss — Modern Replacement for netstat

ss is faster and more powerful than netstat.

Command:

ss -tulnp

Example output:

LISTEN 0 128 0.0.0.0:80 users:(("nginx",pid=1023))

DevOps Use Case

Check if application is listening on port:

ss -tulnp | grep 8080

Output:

LISTEN 0 128 0.0.0.0:8080 java

Means Java service running on port 8080.


6️⃣ traceroute — Trace Network Path

What it does

traceroute shows the path packets take through routers to reach a destination.

Command:

traceroute google.com

Example output:

1 192.168.1.1
2 10.10.1.1
3 172.217.169.14

Each line is a router hop.


DevOps Use Case

Example problem:

Server reachable but slow

Run:

traceroute api.example.com

If a hop has high latency:

200 ms

Network routing problem exists.


7️⃣ dig — DNS Lookup Tool

dig queries DNS servers.

Command:

dig google.com

Example output:

google.com. 300 IN A 142.250.183.206

Meaning:

Field Meaning
google.com domain
A record type
IP resolved address

DevOps Use Cases

Verify DNS resolution

dig api.company.com

Query specific DNS server

dig @8.8.8.8 google.com

This checks Google DNS directly.


8️⃣ nslookup — DNS Query Tool

nslookup is another tool to query DNS.

Command:

nslookup google.com

Output:

Name: google.com
Address: 142.250.183.206

DevOps Use Case

Check DNS misconfiguration.

Example problem:

API domain not resolving

Run:

nslookup api.example.com

If result:

server can't find api.example.com

DNS record missing.


🚀 Real DevOps Debugging Scenario

Problem

Users cannot access website.


Step 1 — Check connectivity

ping example.com

Step 2 — Verify DNS

dig example.com

Step 3 — Check service

ss -tulnp | grep 80

Step 4 — Test HTTP response

curl http://example.com

Step 5 — Check network route

traceroute example.com

This workflow helps isolate the problem quickly.


📊 Summary

Command Purpose
ping test connectivity
curl test HTTP/API
wget download files
netstat view network connections
ss modern netstat
traceroute trace network path
dig DNS lookup
nslookup DNS query

⭐ Why These Tools Matter in DevOps

DevOps engineers use these commands to:

  • debug microservice communication

  • verify API availability

  • diagnose DNS issues

  • check open ports

  • analyze network routes

  • validate infrastructure connectivity

These tools are essential for troubleshooting distributed systems and cloud infrastructure.


A real production incident often involves multiple services failing to communicate with each other. In a microservices architecture, one service depends on several others (API, database, authentication service, etc.). If connectivity breaks anywhere, the entire application may fail.

Below is a step-by-step example of how DevOps engineers debug a microservice connectivity issue in production.


🚨 Real Production Incident: Debugging a Microservice Connectivity Issue

Scenario

A company runs a web application with this architecture:

User → Nginx → API Service → Database

Suddenly users start reporting:

500 Internal Server Error

The website is not working.

The DevOps engineer needs to identify where the failure occurs.

Step 1️⃣ Verify the Website Response

First, check if the service responds.

Command:

curl http://example.com

Response:

HTTP/1.1 500 Internal Server Error

This confirms the server is reachable but something inside the system is failing.


Step 2️⃣ Check Nginx Logs

Logs usually show the first hint.

tail -f /var/log/nginx/error.log

Example error:

connect() failed (111: Connection refused) while connecting to upstream

This means:

Nginx → cannot connect to backend API

So the problem is likely with the API service.


Step 3️⃣ Check If the API Service Is Running

Check if the service port is listening.

ss -tulnp | grep 8080

Example output:

LISTEN 0 128 0.0.0.0:8080 users:(("java",pid=2431))

If no output appears:

API service is not running

Restart service:

systemctl restart api-service

If service is running, move to next step.


Step 4️⃣ Test API Directly

Now test the backend API.

curl http://localhost:8080/health

Example output:

{"status":"DOWN"}

This indicates the API itself has an internal problem.


Step 5️⃣ Check API Logs

Look at application logs.

tail -f /var/log/app.log

Example output:

Database connection failed

Now we know the issue is between:

API → Database

Step 6️⃣ Verify Database Connectivity

Test database port.

Example MySQL:

ss -tulnp | grep 3306

If nothing appears:

Database service stopped

Start database:

systemctl start mysql

Step 7️⃣ Check DNS Resolution

If the API connects using a hostname:

db.internal.company

Check DNS:

dig db.internal.company

Example failure:

NXDOMAIN

DNS record missing.


Step 8️⃣ Check Network Connectivity

Test connectivity between services.

ping db.internal.company

Or test port access:

nc -zv db.internal.company 3306

Example result:

Connection refused

Database may be blocked by firewall.


Step 9️⃣ Check Firewall Rules

Example:

sudo ufw status

Or cloud firewall rules (AWS security groups).

If port is blocked:

Allow port 3306

Step 🔟 Confirm System Is Fixed

Restart services:

systemctl restart nginx
systemctl restart api-service

Test again:

curl http://example.com

Output:

HTTP/1.1 200 OK

System working again.


🔍 Root Cause Analysis

Final root cause:

Database service crashed → API could not connect → Nginx returned 500 error

📊 Typical DevOps Debugging Flow

DevOps engineers usually follow this troubleshooting chain:

  1. Client → Website response

  2. Web server logs

  3. Backend service health

  4. Application logs

  5. Database connectivity

  6. DNS resolution

  7. Network routes

  8. Firewall rules

Each step narrows down the exact failure point.


🚀 Tools Used in This Investigation

Tool Purpose
curl test HTTP service
tail monitor logs
ss check open ports
dig DNS lookup
ping test connectivity
systemctl manage services

⭐ Key DevOps Lesson

In distributed systems:

Most outages are not code bugs.
They are connectivity problems between services.

Common causes include:

  • crashed service

  • wrong port configuration

  • DNS misconfiguration

  • firewall blocking traffic

  • database failure

Understanding Linux networking and log analysis tool