Open to Work · Germany · Remote / HybridOffen für neue Stellen · Deutschland · Remote / Hybrid
DevOps · SRE · Self-Healing Systems

Muhammad Baqir

DevOps / SRE Engineer focused on infrastructure reliability, automated incident management, self-healing systems, and AI-powered operations.DevOps / SRE Engineer mit Fokus auf Infrastruktur-Zuverlässigkeit, automatisiertes Incident-Management, Self-Healing-Systeme und AI-gestützte Operations.

Muhammad Baqir
AboutÜber mich

I engineer reliable, self-healing infrastructure and automation workflows designed for high availability, rapid incident recovery, and intelligent operations.Ich entwickle zuverlässige, selbstheilende Infrastruktur und Automatisierungsworkflows für hohe Verfügbarkeit, schnelle Incident-Wiederherstellung und intelligente Operations.

My technical focus spans AWS/Azure cloud deployments, Kubernetes container orchestration, Prometheus/Grafana monitoring, and AI-assisted incident management. I specialize in reducing operational overhead, automating incident response, and building self-healing systems that detect, diagnose, and remediate issues without manual intervention.Mein technischer Schwerpunkt liegt auf AWS/Azure-Cloud-Deployments, Kubernetes-Orchestrierung, Prometheus/Grafana-Monitoring und AI-gestütztem Incident-Management. Ich spezialisiere mich auf die Reduzierung von Betriebsaufwand, die Automatisierung von Incident-Response und den Aufbau von Self-Healing-Systemen.

Currently completing my M.Sc. in International Software Systems Science at Otto-Friedrich-Universität Bamberg. I am actively seeking full-time DevOps / SRE / Platform Engineering roles, Werkstudent positions, and Master's thesis opportunities.Derzeit absolviere ich meinen Master in International Software Systems Science an der Otto-Friedrich-Universität Bamberg. Ich suche nach Vollzeitstellen im Bereich DevOps / SRE / Platform Engineering, Werkstudentenstellen oder Masterarbeiten.

−35%downtime reducedAusfallzeit reduziert
−18%cloud cost savedCloud-Kosten gespart
20+systems monitoredüberwachte Systeme
ExperienceBerufserfahrung
Parkyeri — DevOps Engineer
Jun 2022 – Dec 2023
  • Engineered predictive time-series forecasting models using Python and TensorFlow across 200+ inventory items on AWS and Azure, reducing shortages by 20%.Entwicklung von Zeitreihen-Prognosemodellen mit Python und TensorFlow für über 200 Inventarartikel auf AWS und Azure, was Fehlbestände um 20% reduzierte.
  • Supported hybrid cloud migration to AWS and established active infrastructure monitoring via Nagios, Zabbix, and VMware, cutting downtime by 35%.Unterstützung der Hybrid-Cloud-Migration zu AWS und Aufbau von Infrastruktur-Monitoring über Nagios, Zabbix und VMware, was Ausfallzeiten um 35% senkte.
  • Optimized resource utilization and workflow efficiency on cloud environments, reducing cloud operational costs by 18%.Optimierung der Ressourcenauslastung und Effizienz in Cloud-Umgebungen zur Senkung der Cloud-Betriebskosten um 18%.
  • Standardized L1/L2 escalation procedures for a 5-person IT support team, elevating First-Call Resolution (FCR) from 61% to 84%.Standardisierung von L1/L2-Eskalationsprozessen für ein 5-köpfiges IT-Support-Team, wodurch die Erstlösungsrate (FCR) von 61% auf 84% stieg.
Hexagon Helix — DevOps Engineer
Mar 2018 – Mar 2022
  • Centralized metric telemetry and health monitoring across 20+ distributed systems using Prometheus, Grafana, and Azure Monitor.Zentralisierung von Metrik-Telemetrie und System-Monitoring über 20+ verteilte Systeme hinweg mit Prometheus, Grafana und Azure Monitor.
  • Reduced Mean Time to Recovery (MTTR) by 25% by standardizing incident response playbooks and root-cause analysis (RCA) documentation.Reduzierung der durchschnittlichen Wiederherstellungszeit (MTTR) um 25% durch Standardisierung von Incident-Playbooks und Root-Cause-Analysen.
  • Maintained enterprise network infrastructure encompassing DNS, IPsec VPNs, Next-Gen Firewalls, TLS, and Cisco/Juniper routing across 15+ nodes.Betreuung von Enterprise-Netzwerkinfrastruktur mit DNS, IPsec-VPNs, Firewalls, TLS und Cisco/Juniper-Routing über 15+ Knoten.
Work Case StudiesPraxisbeispiele

Case studies from production infrastructure work. Praxisbeispiele aus der Arbeit an Produktionsinfrastruktur.

Parkyeri · DevOps Engineer
AWS Infrastructure Migration & Monitoring

Executed hybrid infrastructure migration to AWS while deploying Nagios, Zabbix, and VMware integration for full-stack telemetry and proactive alert management.Durchführung einer Hybrid-Infrastrukturmigration zu AWS mit Anbindung von Nagios, Zabbix und VMware für Full-Stack-Telemetrie und Alarmierung.

AWS · VMware · Nagios · Zabbix · Telemetry · Hybrid Cloud
−35%infrastructure downtimeInfrastruktur-Ausfallzeit
Parkyeri · DevOps Engineer
Inventory Forecasting & ML Analytics

Implemented machine learning models in Python and TensorFlow on AWS/Azure to forecast inventory demand dynamics across 200+ tracked SKU lines.Implementierung von Machine-Learning-Modellen in Python und TensorFlow auf AWS/Azure zur Nachfrageprognose über 200+ SKU-Linien.

AWS · Azure · Python · TensorFlow · Time-Series · ML Pipelines
−20%inventory shortagesFehlbestände
Parkyeri · DevOps Engineer
IT Service Desk & Escalation Optimization

Restructured incident handling workflows and technical escalation playbooks for a 5-engineer team, raising first-contact resolution rates.Restrukturierung von Incident-Workflows und Eskalations-Playbooks für ein 5-köpfiges Support-Team zur Steigerung der Erstlösungsrate.

ITSM · Incident Handling · L1/L2 Escalation · Troubleshooting
61% → 84%first-call resolutionFirst-Call-Resolution
Hexagon Helix · DevOps Engineer
Centralized Infrastructure Observability

Aggregated metric data across 20+ distributed nodes using Prometheus, Grafana, and Azure Monitor to build real-time health inspection dashboards.Aggregierung von Metrikdaten über 20+ verteilte Knoten mit Prometheus, Grafana und Azure Monitor für Echtzeit-Dashboards.

Prometheus · Grafana · Azure Monitor · Metric Exporters · Observability
20+nodes monitoredüberwachte Knoten
Hexagon Helix · DevOps Engineer
Incident Response & MTTR Reduction

Standardized triage protocols, automated log collection, and published Root Cause Analysis (RCA) documentation to speed up recovery.Standardisierung von Triage-Protokollen, automatische Log-Sammlung und detaillierte Root-Cause-Analysen zur Beschleunigung der Fehlerbehebung.

Incident Response · MTTR Reduction · Triage · Root Cause Analysis
−25%MTTR troubleshooting timeMTTR Fehlerbehebungszeit
Hexagon Helix · DevOps Engineer
Enterprise Network & Security Engineering

Configured and governed secure network infrastructure across 15+ operational endpoints, deploying Cisco/Juniper routing, IPsec VPNs, firewalls, and TLS.Einrichtung und Betreuung von sicherer Netzwerkinfrastruktur für 15+ Endpunkte mit Cisco/Juniper-Routing, IPsec-VPNs, Firewalls und TLS.

Cisco · Juniper · IPsec VPN · Firewalls · TLS · DNS Engineering
15+managed network nodesbetreute Netzwerkknoten
Personal ProjectsEigene Projekte
Self-Healing Infrastructure · AI/RAG · SRE
OpsGuard — Infrastructure Reliability & Self-Healing Platform

AI-powered platform implementing a complete 9-state incident lifecycle with evidence collection, deterministic diagnosis, optional AI/RAG-enhanced root cause analysis, allowlisted remediation with risk-based approval workflow, and automatic recovery verification. Features HTTP health and Prometheus detectors, ChromaDB runbook retrieval, SLO/SLI monitoring (MTTD, MTTR), cost optimization, and rate limiting.AI-gestützte Plattform mit vollständigem 9-State-Incident-Lifecycle, Evidence Collection, deterministischer Diagnose, optionaler AI/RAG-gestützter Root-Cause-Analyse, Allowlisted Remediation mit Risikobasierter Approval-Workflow und automatischer Recovery-Verification. Mit HTTP-Health- und Prometheus-Detektoren, ChromaDB-Runbook-Retrieval, SLO/SLI-Monitoring und Cost-Optimierung.

Python 3.11 · FastAPI · PostgreSQL · Redis · ChromaDB · Prometheus · Grafana · Docker · Kubernetes · Helm · Terraform · Argo CD · GitHub Actions
Customer Churn Analytics Service

End-to-end MLOps service for customer churn inference, feature drift monitoring, and explainability using Streamlit, FastAPI, Scikit-Learn, and Evidently AI.End-to-End MLOps-Service für Churn-Inferenz, Daten-Drift-Monitoring und Erklärbarkeit mit Streamlit, FastAPI, Scikit-Learn und Evidently AI.

Python · Scikit-Learn · FastAPI · Streamlit · Evidently AI · Docker
Real-Time Data Engineering Pipeline

Scalable ETL data pipeline for live data ingestion, validation, and automated PostgreSQL persistence with FastAPI REST query interfaces.Skalierbare ETL-Datenpipeline zur Ingestion, Validierung und PostgreSQL-Speicherung mit FastAPI REST-Schnittstellen.

Python · Pandas · SQLAlchemy · PostgreSQL · FastAPI
Technical SkillsTechnische Kenntnisse
AWSAzureKubernetesDockerTerraformHelmArgo CDCI/CDLinuxPythonTensorFlowFastAPIPostgreSQLRedisChromaDBPrometheusGrafanaLokiAzure MonitorNagiosZabbixVMwareGitGitHub ActionsTrivyGitleaksCiscoJuniperDNSIPsec VPNFirewallsTLSMLOpsRAGRCASLO/SLIMTTD/MTTRIncident ManagementSelf-Healing Systems
Let's ConnectKontakt

I am open to full-time roles, Werkstudent positions, and Master's thesis projects in DevOps, SRE, Cloud Infrastructure, and MLOps across Germany.Ich stehe für Vollzeitstellen, Werkstudentenstellen und Masterarbeiten in den Bereichen DevOps, SRE, Cloud Engineering und MLOps in Deutschland zur Verfügung.

Get in TouchKontakt Aufnehmen