Job Monitor / Oil & Gas

Digital Technology Senior Specialist – Observability & AI Ops

Baker Hughes · India · Posted 1 hour ago

Company Baker Hughes Location IN-Maharashtra-Pune-7th Floor, Tower 1 & 2, Phoenix Millenium Towers S No. 132, 23. Pune-Bangalore Highway, India Employment type Full-time Posted 1 hour ago Listed via Baker Hughes
DT Senior Specialist – Observability & AI Ops Would you like to help shape and implement our Digital Technology teams' strategic direction? Are you passionate about helping improve observability and digital operations? Join our Digital Technology team! We operate at the heart of Baker Hughes digital transformation journey. Our team delivers enterprise observability and AIOps capabilities that help technology teams detect issues earlier, troubleshoot faster, and improve service performance across cloud, infrastructure, and application environments. Partner with the best As an AIOps & Observability Engineer, you will support the implementation, enhancement, onboarding, and day-to-day operations of observability platforms, with a focus on Elastic Stack capabilities and practical SRE-driven operational outcomes. As a Senior AI Ops Engineer, you will be responsible for: Implement and manage the life cycle of enterprise observability platform based on elastic tech stacks [not limited to] Kibana, Logstash, Beats, Elastic Agent, Fleet, Elastic APM components, etc. Onboard full suite of 20,000 plus devices under observability umbrella Onboarding infrastructure, cloud services, applications, and platforms into observability and monitoring solutions. Building and maintaining dashboards, visualizations, alerts, and operational reports for technology teams. Configuring log, metric, trace, uptime, and APM data collection across supported environments. Assisting with data ingestion, parsing, enrichment, and retention activities. Supporting incident investigation, troubleshooting, and root cause analysis using observability data. Collaborating with cloud, infrastructure, application, and SRE teams to improve system reliability and service visibility. Contributing to automation initiatives using scripting, Infrastructure as Code, and repeatable deployment practices. Participating in observability platform upgrades, patching, performance tuning, and operational support activities. Creating and maintaining runbooks, knowledge articles, dashboards standards, and operational runbooks related documentation. Contributing to continuous improvement of observability practices, monitoring coverage, and operational readiness. Fuel your passion Have 7+ years – SRE/DevOps experience in enterprise-scale or mission-critical environments Have 5+ years – Cloud / Application / Platform operations and administration (AWS, Azure, hybrid or multi-cloud) Have 5+ years – Automation, CI/CD, and scripting proficiency (Python, Bash, PowerShell, Ruby, or equivalent) Have 5+ years - Exposure to containers and cloud-native platforms such as Docker, Kubernetes, Prometheus, or Grafana. Have 3+ years – Proven experience administering Elastic Observability platforms across the full lifecycle, including deployment, maintenance, upgrades, patching, and capacity scaling. Preferred qualifications AWS or Azure Associate-level certification, or equivalent practical cloud operations experience. Elastic Certified Engineer or equivalent observability platform certification Familiarity with infrastructure as code (GitHub Actions, CloudFormation, Terraform, Ansible) for repeatable automation Exposure to cloud-native observability frameworks (Open Telemetry, service meshes) Experience documenting runbooks, playbooks, and consumption guides for SMEs Process knowledge such as Agile and/or ITIL Must have Technical Skills Strong background in observability platforms (Elastic.io stack preferred: Elasticsearch, Kibana, Logstash, Beats, Elastic APM, and Fleet/Elastic Agent) Telemetry Fundamentals: Strong, practical understanding of fundamental observability concepts, including the collection and analysis of logs, metrics, traces, and synthetic monitoring. Experience with administration of leading observability platforms (Grafana, Graylog, Splunk, Sumo Logic, Tanzu, or open-source equivalents) including lifecycle management (Kubernetes, Docker, Prometheus, and Grafana installation, patching, upgrades, scaling). Strong knowledge of distributed infrastructure domains (network, servers, VMs, AWS, Azure, databases) from an observability perspective. Proven ability to design and tune scalable ingestion pipelines for diverse, globally distributed data sources. Flexibility to adapt and evolve observability standards per domain needs while ensuring consistency across the enterprise. Operational ownership mindset — accountable for uptime, reliability, and lifecycle management of the hosted observability platform. Incident management and SRE practices: monitoring, alerting, troubleshooting, root cause analysis, and postmortems. Proficiency in automation and scripting (Python, Bash, PowerShell, Ruby, etc.) for ingestion, upgrades, and operational tasks. Familiarity with REST APIs and tools like Postman, plus DevOps constructs (GitHub, Jenkins, CI/CD pipelines, serverless technologies). Configuration Proficiency: Demonstrated proficiency in managing system configurations using YAML-based
Explore more
633 Baker Hughes jobs → 244 jobs in India →
Oil & Gas Alert tracks 21+ employer career pages and delivers daily digests of new vacancies. Set your own keywords →