Showing "DevOps Engineer" roles at UnitedHealth Group
9 of 10,000+ jobs
Associate Software Engineering Manager
UnitedHealth Group•Noida, Uttar Pradesh, IN
5+ Years Exp
Posted: 10/8/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
• Design, develop, and integrate AI-powered features and solutions into enterprise applications to enhance user experience and business outcomes
• Collaborate with SMES, ML engineers, and product teams to operationalize machine learning models and AI services in production environments
• Build and maintain scalable APIs and services that leverage Generative AI, Large Language Models (LLMs), and intelligent automation capabilities
• Evaluate, fine-tune, and integrate AI models and frameworks to address business use cases and improve application functionality
• Implement responsible AI practices, ensuring solutions meet security, privacy, compliance, and ethical AI standards
• Develop and optimize Retrieval-Augmented Generation (RAG) solutions, vector databases, and knowledge-grounding capabilities
• Monitor AI solution performance, model quality, accuracy, and operational reliability in production environments
• Stay current with emerging AI technologies, industry trends, and best practices, and recommend their adoption where appropriate
• Partner with stakeholders to identify opportunities for leveraging AI to improve productivity, automation, and decision-making
• Contribute to AI governance, model evaluation frameworks, prompt engineering practices, and AI solution architecture standards
• Mentor team members on AI-enabled development practices and promote the adoption of modern AI engineering approaches
• Design, develop, test, and maintain scalable, reliable, and secure software applications
• Lead the end-to-end delivery of complex features and projects, from requirements gathering through deployment and support
• Collaborate with product managers, architects, designers, and cross-functional engineering teams to define and implement technical solutions
• Write clean, maintainable, and well-documented code following industry best practices and coding standards
• Conduct code reviews, provide constructive feedback, and mentor junior engineers to promote engineering excellence
• Troubleshoot and resolve complex technical issues, ensuring optimal application performance and reliability
• Drive system architecture discussions and contribute to technical design decisions
• Improve software quality through automated testing, CI/CD pipelines, and observability practices
• Participate in Agile development processes, including sprint planning, estimation, stand-ups, and retrospectives
• Continuously evaluate and adopt new technologies, tools, and development practices to improve productivity and product quality
• Collaborate with stakeholders to understand business requirements and translate them into technical solutions
• Ensure compliance with security, privacy, and operational standards throughout the software development lifecycle
• Support production systems, analyze root causes, and implement preventive measures for recurring issues
• Contribute to technical documentation, knowledge sharing, and process improvements across the engineering organization
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
• Bachelor's or Master's degree in Computer Science, Engineering, or a related field
• 5+ years of professional software development experience
• Hands-on experience with DevOps practices, CI/CD, and containerization technologies
• Hands-on experience with Generative AI technologies, including Large Language Models (LLMs) such as OpenAI, Azure OpenAI, Anthropic Claude, or similar platforms
• Experience with cloud platforms (Azure, AWS, or Google Cloud)
• Experience mentoring engineers and leading technical initiatives
• Experience building applications using AI services such as Azure AI, Azure OpenAI, Copilot technologies, or equivalent cloud-based AI platforms
• Experience developing and deploying AI/ML-powered applications in enterprise environments
• Experience building AI-powered solutions using Retrieval-Augmented Generation (RAG), vector databases, semantic search, and knowledge-grounding techniques
• Experience with AI Agent development and multi-agent systems
• Experience integrating AI capabilities into web, mobile, or enterprise applications through APIs and microservices architectures
• Knowledge of MLOps practices, model deployment, monitoring, and lifecycle management
• Understanding of Responsible AI principles, including security, privacy, governance, fairness, and compliance considerations
• Solid understanding of machine learning fundamentals, natural language processing (NLP), and intelligent automation concepts
• Solid understanding of microservices architecture, APIs, databases, and distributed systems
• Familiarity with AI orchestration frameworks such as LangChain, Semantic Kernel, LangGraph, or equivalent technologies
• Familiarity with Azure AI Services, Azure AI Foundry, Azure Machine Learning, Copilot technologies, or comparable cloud AI platforms
• Proficiency in prompt engineering, AI model evaluation, and optimization techniques
• Solid proficiency in one or more programming languages such as Java, C#, Python, JavaScript, or TypeScript
• Demonstrated ability to evaluate emerging AI technologies and apply them to solve business problems effectively
• Proven excellent problem-solving, communication, and collaboration skills
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Senior DevOps Engineer
UnitedHealth Group•Gurugram, Haryana, IN
4+ Years Exp
Posted: 9/8/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
• Maintain and administer cloud environments across platforms; provide support for a broad range of cloud services, ensuring consistent support delivery and timely resolution of issues
• Implement and enforce governance policies, such as access control, tagging, cost management, guardrails, SCPs, and resource monitoring, to ensure adherence to company-wide security standards
• Proactively identify security vulnerabilities and mitigate risks
• Implement, and manage automation scripts and tools to streamline cloud operations, including infrastructure provisioning, configuration, and monitoring
• Design, implement, and manage CI/CD pipelines for the automated deployment of applications and services in a scalable and secure manner
• Monitor system performance, analyze metrics, and implement improvements to ensure efficient utilization of cloud resources
• Collaborate with peers on troubleshooting technical issues; assist in implementing technical strategies established by the team
• Aid in implementing disaster recovery strategies to ensure data protection and maintain business continuity in accordance with the team's guidelines
• Act with integrity, professionalism, and personal responsibility to uphold Optum's respectful and courteous work environment
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
• 4+ years of recent experience in Cloud IT Operations
• Solid understanding of cloud platforms such as AWS, Azure, or GCP
• Proficiency in data analysis and programming languages like Python, PowerShell
• Proficiency in Infrastructure as Code (Terraform) and Advanced CI/CD expertise with GitHub actions
• Proficiency with Cloud monitoring and Cloud networking
• Proven ability to utilize AI for automation and analysis
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Software Engineering Lead Service Now Development
UnitedHealth Group•Hyderabad, Telangana, IN
0+ Years Exp
Posted: 5/8/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Software Engineering Lead: Service Now Development, JavaScript, Server - Side Scripting, ITSM
Primary Responsibilities:
• Develop and maintain ServiceNow applications, business rules, script includes, workflows, and platform integrations
• Support business areas onboarding and deliver enhancement requests aligned with business requirements
• Design and implement AI-driven and automation solutions to improve operational efficiency, productivity, and user experience
• Participate in architecture reviews, code reviews, technical design sessions, and solution governance activities
• Support ServiceNow platform upgrades, security remediation efforts, vulnerability management, and compliance initiatives
• Collaborate closely with Product Owners, QA teams, Architects, Operations teams, and other stakeholders to ensure successful solution delivery
• Provide production support, incident troubleshooting, root cause analysis, and timely issue resolution
• Contribute to SaaS migration, cloud modernization, and platform transformation initiatives
• Ensure adherence to best practices, coding standards, performance optimization, and platform governance requirements
• Support Agile delivery processes, including sprint planning, backlog refinement, development, testing, and release activities
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
• Bachelor's degree or equivalent experience
• Experience in github, aware of CICD and code management
• Proven excellent written and verbal communication skills, co-ordination skills
Preferred Qualifications:
• Healthcare Domains exposure
• AI exposure, able to use LLM models etc.
• Airflow exposure
• Experienced ServiceNow professional responsible for designing, developing, administering, and optimizing enterprise ServiceNow solutions
• Experience to partner with cross-functional teams to deliver scalable applications, integrations, automation capabilities, and platform enhancements while ensuring security, performance, operational excellence, and alignment with business objectives
Key Technical Skills:
• ServiceNow Development and Administration
• ServiceNow ITSM Platform
• JavaScript and Server-Side Scripting
• Flow Designer
• REST and SOAP Integrations
• Performance Tuning and Troubleshooting
• Agile/Scrum Methodologies
• DevOps and CI/CD Practices
• Platform Security and Vulnerability Management
• AI and Automation Solutions
• Cloud and SaaS Migration Support
• Technical Design and Architecture Review
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Software Engineer Java Full Stack
UnitedHealth Group•Hyderabad, Telangana, IN
0+ Years Exp
Posted: 3/8/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
• Develop and maintain full‑stack web applications using Angular, Java, and SQL
• Build secure, scalable APIs and microservices, including event‑driven services using Kafka
• Implement and manage CI/CD pipelines and automate deployments using GitHub and DevOps tools
• Containerize and deploy applications on Kubernetes, ensuring scalability and reliability
• Collaborate across teams to deliver high‑quality solutions and resolve production issues
• Follow best practices for security, performance, and code quality
• Take ownership of features from design through production support
• Continuously learn and adopt new technologies, including AI‑related tools and capabilities
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regard to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
• Bachelor's degree in Computer Science, Engineering, or equivalent experience
• Solid experience with Angular, Java, SQL, and RESTful services
• Hands‑on experience with microservices, Kafka, and event‑driven systems
• Experience with GitHub, CI/CD pipelines, and DevOps practices
• Working knowledge of Docker and Kubernetes
• Proven solid problem‑solving skills and ability to work across front‑end, back‑end, and cloud environments
• Openness to continuous learning; AI builder or AI development exposure is a plus
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone - of every race, gender, sexuality, age, location and income - deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
DevOps Engineer Azure
UnitedHealth Group•Gurugram, Haryana, IN
5+ Years Exp
Posted: 5/7/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
• Build and maintain Azure resources using modern Infrastructure-as-Code tooling (Terraform, Ansible)
• Build and maintain pipelines and automation through GitOps (GitHub Actions, JPAC, etc.)
• Create and own platform level services on top of Kubernetes (AKS, Azure Container Apps)
• Develop automation scripts by submitting PRs and participating in code reviews (Github)
• Use AI tools like GitHub Copilot and ChatGPT to enhance productivity, automation, and code quality
• Learn and apply new AI tools and techniques to improve development workflows and solutions
• Stay up-to-date with the latest advancements in AI and integrate them into daily tasks and projects
• Build monitoring and alerting templates for various cloud metrics (Splunk, Dynatrace, Azure app insights)
• Participate in a shared on-call rotation
• Works with less structured, more complex issues
• Analyzes and investigates
• Provides explanations and interpretations within area of expertise
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
• Graduate degree or equivalent experience
• 5+ years of experience defining, designing, and implementing CI/CD systems (such as GitHub Actions, Jenkins, etc.)
• 3+ years of experience building and troubleshooting cloud networks
• 2+ years of experience with Terraform
• Experience with Monitoring tools such as Dynatrace, Splunk, Grafana, and Prometheus
• Experience with AI tools and integration (like ChatGPT, M365 Copilot GitHub Copilot, etc.)
• Experience with advanced scripting
• Knowledge of AI model integration into applications
• Solid knowledge of Azure networking, such as configuring virtual networks, firewalls, load balancers, and VPNs
• Expertise in identifying and mitigating infrastructure security vulnerabilities
• Expertise in managing cost optimized infrastructure setups
• Expertise in setting monitors, triggers, alerts and managing the infra setup on highest availability and reliability standards
• Proven team player attitude and willingness to collaborate with others
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Software Engineer Dot Net Fse
Unitedhealth Group•Pune, Maharashtra, IN
0+ Years Exp
Posted: 25/6/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
• Staying plugged into emerging technologies/industry trends and applying them into operations and activities
• Open-Source Tools and Framework
• Large and more complex components, while influencing overall Data architecture and patterns
• Influences the team designs and solutions Mentors software engineers through code reviews, and hands-on design sessions
• Code reviews are appropriately broken down in to reviewable chunks
• Contributes meaningfully to code reviews of teams work, providing collaborative guidance in areas of strength
• Continues to receive guidance in own code reviews primarily around solution refinement, rather than overall direction Able to explain implementation decisions and push back appropriately
• Learns from feedback and applies to future deliverables
• Solves more complex problems API's / Data Structures/ Data Models/Algorithms/ Application Sequences are thoughtfully designed Solutions are well integrated, testable, maintainable and performant
• Appropriately leverages existing solutions and adapts for reuse
• Delivers solutions with the appropriate toolset (languages, algorithms, patterns and frameworks) for the constraints and conditions of the business, team and product
• Choose refactor opportunities to drive down tech debt, in alignment with sprint and program goals
• Incorporates automation in testing, build, and deployment processes to drive team efficiencies
• Introduces automation to replace repeated manual processes demonstrating measurable improvement
• Work with internal stakeholders, get their priorities, build alignment on prioritization and synthesize them into quarterly roadmap for the platform
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Builder Responsibilities:
Design, develop, and deploy AI-powered solutions using no-code, low-code, and advanced platforms, translating business needs into scalable applications that enhance products, workflows, and decision-making.
Required Qualifications:
• Bachelor's degree in Computer science or related field
• Solid experience in development using C#, .Net Core Web API and/or Java and Angular and/or React
• Experience with SQL/NoSQL databases
• Solid experience in DevOps best practices in cloud native environment
• Experience in software development
• Experience delivering solutions with a focus on C#, .Net Core Web API, Angular/React along with Python - AI/ML exp and automation
• Experience with the technologies for this Project:
• Java
• Angular / React, Databricks, Python for AI/ML and automation
• Open-Source Tools and Framework
• Azure Public Cloud
• Git Actions
• Git
• Sonar
• Good knowledge of Entity Framework Core
• Very good understanding of agile software development practices and role of product management in the same
• Solid understanding of product management principles - prioritization, preparing roadmaps, cross-functional team collaboration
• Very good high-level understanding of building software platforms
• Healthcare knowledge
• Proven excellent problem-solving skills, understanding of Data Structures
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Principal Data Scientist
UnitedHealth Group•Gurugram, Haryana, IN
10+ Years Exp
Posted: 24/6/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
• Provide technical leadership and mentorship to data scientists, fostering innovation and excellence across teams
• Lead end-to-end delivery of complex technical and AI/ML-driven initiatives, ensuring alignment with business objectives, timelines, budget, and quality standards
• Own and prioritize the product backlog, translating business needs into user stories, including AI/ML use cases with clear success criteria
• Drive AI/ML adoption across products and platforms, ensuring integration into workflows, scalable deployment, and alignment with enterprise strategy
• Establish real-time measurement frameworks (KPIs, ML metrics, adoption rates, model performance, business impact) using dashboards and analytics tools
• Partner with business, data science, engineering, and architecture teams to define product vision, roadmap, and value-driven prioritization
• Ensure seamless collaboration across engineering, other Tech partners, data, DevOps, and operations for secure, scalable, and compliant solution delivery
• Drive adoption of modern technologies including Java, Python, cloud platforms (Azure/GCP), MLOps
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
• B Tech/ M Tech or MCA
• 10+ years of experience in agile delivery, product lifecycle, or related technical roles
• Experience with Rally / Azure DevOps / Confluence
• Experience of working in Java, Python, Azure Cloud
• Solid understanding of AI DLC, Agile/Scrum, and product delivery
• Technical understanding of systems, APIs, integrations, and enterprise solutions
• Proven excellent communication, facilitation, and cross-functional leadership skills
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Principal Site Reliability Engineer
UnitedHealth Group•Noida, Uttar Pradesh, IN
15+ Years Exp
Posted: 21/6/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
• Define and own the SRE, AI-enabled operations, and observability strategy for the assigned portfolio, aligned with organizational goals and focused on improving reliability, stability, security, scalability, supportability, resilience, automation, and operational excellence across all digital properties
• Act as a senior technical leader who bridges Site Reliability Engineering, software engineering, IT operations, AI engineering, observability platform engineering, cloud/platform teams, data engineering, and business technology leadership
• Provide technical leadership, mentorship, and strategic guidance to senior and mid-level SREs, platform engineers, AI implementation teams, observability engineers, and cross-functional technology teams
• Foster a culture of engineering excellence, continuous learning, proactive reliability, automation-first operations, production ownership, and operational discipline
• Collaborate with engineering, security, architecture, product, data platform, cloud, infrastructure, operations, and business leaders to integrate reliability, observability, AI-enabled automation, and operational best practices into products and platforms from design through deployment
• Report to senior stakeholders and CIO-level leaders on critical paths, operational risks, reliability posture, production readiness, mitigation plans, technology debt, AI adoption opportunities, and strategic SRE initiatives
• Define, govern, and continuously improve enterprise reliability standards, including SLAs, SLIs, SLOs, error budgets, operational risk scoring, production readiness criteria, resilience scorecards, and service health models
• Lead the reliability and peak season readiness initiatives by owning the assessment framework, collaborating with application teams, identifying reliability gaps, and driving critical applications toward 99.999% availability from a resiliency, availability, and reliability perspective
• Architect and govern enterprise-grade monitoring, alerting, and observability standards across lines of business using platforms such as Splunk, Dynatrace, Grafana, DataDog, OpenTelemetry, ServiceNow, cloud-native monitoring tools, and next-generation observability platforms
• Drive the transition from static dashboards and tool-specific monitoring to unified, intelligent, business-impact-aligned observability that provides visibility into application health, infrastructure health, customer experience, service reliability, operational risk, and business impact
• Lead the design and implementation of a modern enterprise observability dashboard and intelligence platform using technologies such as React, JavaScript, TypeScript, REST APIs, Snowflake, Kafka Streaming, cloud-native services, and enterprise data platforms
• Partner with data engineering and platform teams to design scalable data models, telemetry pipelines, event streams, API integrations, and analytical capabilities using Snowflake, relational databases, Kafka, streaming platforms, and observability data sources
• Integrate real-time and near-real-time telemetry from logs, metrics, traces, events, alerts, incidents, change records, service metadata, cloud platforms, infrastructure platforms, and business systems
• Ensure the observability dashboard supports service health views, dependency mapping, role-based views, SLO tracking, alert correlation, incident insights, customer impact analysis, capacity trends, executive reporting, and AI-assisted recommendations
• Lead the strategy, design, and implementation of AI-enabled observability, AIOps, and intelligent automation capabilities to transform incident management from reactive to proactive, predictive, and increasingly autonomous
• Drive implementation of AI and GenAI capabilities for incident triage, impact assessment, log analysis, anomaly detection, event correlation, root cause analysis, knowledge retrieval, runbook recommendation, production readiness validation, and automated remediation
• Partner with engineering and platform teams to integrate LLM-based triage, Agentic AI workflows, AI-powered observability, and automated remediation into SRE workflows, on-call processes, incident response, and operational support models
• Identify practical AI implementation opportunities that reduce alert noise, accelerate root cause analysis, reduce manual toil, improve developer productivity, and deliver measurable improvements in MTTD, MTTA, MTTR, and MTBI
• Work with security, architecture, data governance, and platform teams to ensure AI-enabled solutions are implemented securely, responsibly, explainably, and in alignment with enterprise standards
• Analyze and model system dependencies across applications, APIs, infrastructure, databases, cloud services, message streams, third-party integrations, and business-critical workflows
• Conduct risk and threat modeling for operational scenarios including natural disasters, cloud region failures, cyberattacks, infrastructure failures, software defects, data pipeline failures, dependency failures, and peak-volume business events
• Design and implement resilience patterns such as automated failover, geo-redundancy, circuit breakers, bulkheads, throttling, graceful degradation, blue-green deployments, canary deployments, automated rollback, and self-healing automation
• Lead chaos engineering strategy and execution to proactively identify failure modes, validate system resilience, and improve recovery readiness
• Provide technical leadership across hybrid hosting environments including Unix, Linux, Windows, Azure, AWS, GCP, private cloud, Kubernetes, containers, serverless platforms, and enterprise hosting platforms
• Partner with infrastructure, cloud, network, security, and application teams to ensure platforms are reliable, scalable, secure, observable, resilient, cost-efficient, and supportable
• Lead technology transformation efforts including cloud migration strategy, HCP assessment and adoption, platform modernization, containerization, serverless architecture, open source and inner source adoption, and automation-led operations
• Guide teams on modern technology trends, emerging AI capabilities, evolving observability practices, changing cloud/platform technologies, and new engineering patterns that can improve reliability and operational effectiveness
• Own and drive automation strategy to eliminate manual toil by designing scalable automation frameworks for runbooks, incident response, change validation, operational support, reporting, remediation, and self-service operations
• Define and track toil metrics, automation coverage, operational efficiency metrics, incident trends, reliability improvement outcomes, and continuous improvement opportunities
• Build automation-first operational models using scripting, APIs, workflow automation, CI/CD integration, AI-assisted workflows, and reusable engineering patterns
• Improve operational tooling and frameworks by evaluating, selecting, standardizing, and governing tools across the SRE, observability, AI operations, and platform engineering portfolio
• Act as a senior gatekeeper for production changes by establishing change governance processes, operational risk scoring, AI-assisted readiness validation, rollback validation, and release reliability standards
• Lead incident response for P1 and P2 incidents, including war room facilitation, executive communication, technical triage, impact assessment, recovery coordination, root cause analysis, and post-incident review processes
• Respond to platform emergencies, alerts, and escalations from customer support, business operations, application teams, and technology partners while ensuring root cause is addressed and corrective actions are implemented
• Leverage ServiceNow and ITSM processes for incident, problem, change, knowledge, configuration, and service management at enterprise scale
• Participate in and lead on-call rotation, setting the standard for on-call excellence, operational readiness, knowledge sharing, escalation management, and continuous improvement
• Create and maintain architectural diagrams, flow diagrams, runbooks, operational playbooks, executive-level reports, service health documentation, dashboard documentation, and AI-enabled operational process documentation
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
• Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field preferred.
• 15+ years of overall experience in the IT industry across software development, infrastructure, operations, platform engineering, cloud engineering, production support, or enterprise technology delivery.
• 9+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, Production Engineering, DevOps, Cloud Operations, or a similar role with demonstrated leadership in driving reliability at enterprise scale
• 9+ years of experience designing, implementing, and governing monitoring, alerting, and observability architectures for cloud, hybrid, and enterprise software solutions using tools such as Splunk, Dynatrace, DataDog, Grafana, OpenTelemetry, ServiceNow, or similar platforms
• 7+ years of coding or scripting experience with two or more of the following: Java, Python, Go, JavaScript, TypeScript, C#, C/C++, Perl, PowerShell, Shell scripting, Mainframe technologies, or similar languages
• 5+ years of experience building, designing, integrating, and programmatically consuming REST APIs at scale
• 2+ years of experience mentoring and providing technical leadership to SRE engineers, software engineers, platform engineers, observability engineers, AI engineers, or cross-functional technology teams
• Solid hands-on experience implementing SRE practices across large-scale enterprise applications, including SLAs, SLIs, SLOs, error budgets, monitoring, alerting, incident response, capacity planning, performance engineering, resilience engineering, and production readiness
• Demonstrated experience defining, managing, and operationalizing SLAs, SLIs, SLOs, error budgets, and reliability metrics as operational standards
• Proven practical experience with AI implementation, AIOps, AI-enabled observability, intelligent incident detection, event correlation, anomaly detection, automated response, or LLM-based triage
• Experience identifying AI use cases, designing implementation patterns, integrating AI capabilities into operational workflows, and measuring business or operational outcomes
• Experience building or supporting observability dashboards, operational intelligence platforms, service health portals, or executive reporting solutions
• Experience with modern front-end or dashboard development technologies such as React, JavaScript, TypeScript, HTML, CSS, REST APIs, UI components, and data visualization frameworks
• Experience working with data platforms such as Snowflake, SQL Server, PostgreSQL, MySQL, or similar relational, analytical, or operational data stores
• Experience with streaming or event-driven platforms such as Kafka, Kafka Streams, event hubs, message queues, or similar technologies
• Experience integrating observability data from logs, metrics, traces, events, alerts, incidents, changes, service metadata, infrastructure platforms, cloud platforms, and business systems.
• Experience with automation and deployment tools such as Terraform, Ansible, Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, Argo CD, Helm, Kubernetes operators, or similar tools
• Experience with programmatic interaction with relational databases and data-driven operational decision-making
• Experience leading incident response for P1/P2 production incidents, including war room facilitation, executive stakeholder communication, root cause analysis, and post-incident review processes
• Experience leveraging ServiceNow or similar ITSM platforms for incident, problem, change, knowledge, configuration, and service management processes
• Experience in health care, insurance, financial services, government programs, regulated environments, or large-scale enterprise technology operations
• Solid understanding of hybrid hosting and infrastructure platforms including Unix, Linux, Windows, Azure, AWS, GCP, private cloud, containers, Kubernetes, serverless platforms, and enterprise hosting platforms
• Familiarity with GenAI, Agentic AI, LLM-based assistants, AI copilots, prompt engineering, semantic search, RAG patterns, vector databases, model integration, AI governance, and responsible AI practices
• Proven track record of planning, supporting, or improving 99.999% availability for critical applications in production environments
• Proven solid architectural understanding of engineering fundamentals including unit testing, performance testing, chaos engineering, code reviews, telemetry, Agile, DevOps, CI/CD, security, API design, and production readiness
• Proven deep expertise in CI/CD pipelines, containerization, serverless architecture, public cloud, private cloud, application observability, messaging, streaming architecture, and platform automation
• Demonstrated ability to guide technical priorities, conduct design reviews, influence architecture decisions, define engineering standards, and drive adoption of modern technology practices
• Proven ability to evaluate emerging technologies, understand changing technology dynamics, guide teams on adoption strategy, and translate new technology capabilities into practical enterprise implementation plans
• Proven ability to communicate effectively with technical and non-technical, globally distributed audiences, including presenting to senior leadership and CIO-level stakeholders on reliability posture, AI initiatives, operational risk, and strategic technology direction
• Proven solid technical writing skills, including creating architectural diagrams, flow diagrams, runbooks, end-user documentation, operational playbooks, executive-level reports, and technology strategy documents
• Flexibility to support 24x7 operations through shift-based, on-call, and rotational support models
Preferred Qualification:
• AI Dojo certification Level 1, Level 2, and Level 3
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
Senior Software Engineer Python React
UnitedHealth Group•Hyderabad, Telangana, IN
4+ Years Exp
Posted: 17/6/2026
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
• Develop, test, deploy, maintain and continuously improve software
• Translate product concepts into project commitments that deliver incremental value to our customers frequently and with high quality
• Participate in Agile / Scrum methodology to deliver high-quality software releases
• Performing all phases of software engineering including requirements analysis, application design, code, test, deploy, and support
• Troubleshoot production support issues post-deployment and come up with solutions as required
• Be innovative in solution design and development to meet the needs of the business
• Expected to adapt in dynamic and collaborative work environment and make independent decision
• Collaborate to meet project timelines
• Basic, structured, standard approach to work
• Works with less structured, more complex issues
• Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regard to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
• Bachelor's degree or equivalent experience
• 4+ years of coding experience with 4+ years of the following languages Java, Python or JavaScript with a willingness and ability to learn new ones along with Microservices
• Sound knowledge as AI Builder using any of LLM models
• Good Experience working with continuous integration / continuous delivery tools, REST API development, serverless architecture, containerization, IaC, public / private cloud, application observability and / or messaging / stream architecture
• Experience working within an Agile / Scrum Methodology
• Solid understanding of engineering fundamentals: unit testing, code reviews, Agile and DevOps
Required Qualification:
• Experience in Splunk & Postgress
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone - of every race, gender, sexuality, age, location and income - deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.