JD:
Those are the skill sets we need, hands-on with:
vLLM / LLM inference serving
GPT-OSS / Harmony parser troubleshooting
Tool/function calling debugging
LLM token, context-window and output parsing
Prompt caching / KV-cache optimization
NVIDIA GPU monitoring and performance tuning
Tensor Parallelism (TP2/TP4) tuning
Concurrency, batching and throughput optimization
Python-based AI application debugging
Docker / Kubernetes production operations
API Gateway / reverse proxy troubleshooting
HTTP troubleshooting
TLS / certificate / connection-reset debugging
Load balancer troubleshooting – ALB/NLB
OpenTelemetry / distributed tracing
Langfuse or equivalent LLM observability
Centralized logging / Kibana
End-to-end correlation IDs and request tracing
Retry, timeout, backoff and circuit-breaker design
Fallback / graceful-degradation patterns
Load, stress and soak testing
Latency analysis: P50/P95/P99, TTFT, tokens/sec
Capacity planning and saturation analysis
Production incident management and RCA
CI/CD, configuration versioning and rollback