Lead System Engineer – Turbine Control & Automation Systems
VINCI•AE
10+ Years Exp
Posted: 16/9/2026
Actemium Emirates Projects (AEP), an operational entity of Cegelec Abu Dhabi, is a specialized engineering and contracting business unit focused on electrical, instrumentation, and control systems projects for the Oil & Gas industry. From initial design through to construction and commissioning, AEP delivers end-to-end solutions tailored to the sector’s technical and operational demands.
Active in Abu Dhabi and the Oil & Gas sector since 1980, AEP has successfully executed a wide range of projects, either directly for ADNOC (Abu Dhabi National Oil Company) subsidiaries or as a specialized subcontractor supporting major contractors.
Actemium Emirates Projects benefits from the expertise, resources, and international network of VINCI Energies. Our team is composed of experienced project managers, skilled engineers, and a strong in-house construction workforce — all committed to delivering reliable, agile, and locally adapted solutions, while embracing a dynamic approach to innovation and sustainable performance.
We are seeking an experienced Lead System Engineer to support a major GE Mark VI to Mark VIe Turbine Control System Upgrade project within a leading Oil & Gas facility.
The successful candidate will provide technical leadership throughout engineering, system integration, FAT/SAT, commissioning, shutdown execution, startup, and performance testing activities.
Key Responsibilities
· Lead the migration and upgrade of GE Mark VI to Mark VIe turbine control systems.
· Review and validate control logic, sequencing, protection schemes, alarms, trips, operating philosophies, and cybersecurity compliance.
· Manage interfaces with DCS, ESD, F&G, Compressor Control Systems, Generator Controls, Electrical Systems, and third-party packages.
· Lead design reviews, Factory Acceptance Tests (FAT), Site Acceptance Tests (SAT), commissioning, and startup activities.
· Develop migration, cutover, and startup strategies to minimize plant downtime.
· Provide technical leadership during shutdown execution, system migration, startup, and reliability runs.
· Troubleshoot hardware, software, communication, and field interface issues.
· Ensure compliance with current Oil & Gas cybersecurity requirements and industry standards.
· Act as the primary technical interface between Client, EPC/PMC, OEMs, vendors, and site teams.
Requirements
• Bachelor's Degree in Instrumentation, Control, Electrical, Electronics Engineering.
• Minimum 10 years' experience in Oil & Gas control and automation systems, with overall experience of 15+ years preferred.
• Strong hands-on experience with:
• GE Mark VI / Mark VIe
• ToolboxST / ControlST
• CIMPLICITY
• Turbine Control Systems
• DCS Integration
• Compressor Control Systems
• Industrial Cybersecurity Requirements
• Proven experience in FAT, SAT, site commissioning, startup support, loop checking, troubleshooting, and brownfield upgrade projects.
• Willingness and capability to be mobilized to site during shutdown, commissioning, startup, and performance testing activities.
• Prior Mark VI to Mark VIe migration experience is highly desirable.
• Experience on ADNOC / ADNOC Gas projects will be a significant advantage.
Candidates with demonstrated expertise in GE turbine control systems, DCS and compressor control integration, and hands-on experience supporting commissioning, startup, and shutdown activities in ADNOC or similar Oil & Gas operating facilities are encouraged to apply.
Why join us ?
Our DNA: Trust, Entrepreneurial Spirit, Solidarity, Autonomy, Responsibility
Joining us means becoming part of a large group while enjoying the agility and warmth of a human-sized company!
A personalized onboarding journey from day one, with tailored career follow-up
Opportunities for growth, training, and mobility within a fast-growing international group
Close and supportive management
Recognition of employee performance through a company savings plan
Pride in shared achievements
☘️ Respect for the environment and local communities in the countries where we operate (Human Rights Guide)
Commitment to the health and safety of our employees
At VINCI Energies Oil & Gas, CSR initiatives are deeply embedded in our activities, our ways of working, and ultimately, in our DNA
Machine Learning Engineer, Applied AI, Deployed
Brain Co.•Abu Dhabi, AE
0+ Years Exp
Posted: 16/9/2026
Our Mission
Rebuild how the world works, to make institutions work better for the people they serve.
About Brain Co.
Brain Co. builds AI-native operating systems for large, regulated institutions. Each system is built for a specific industry, powered by agents that push real workflows forward. Underneath it all is Atlas, our proprietary platform that keeps customers in control, secure by design, and never locked into one model.
Why Now
Brain Co. is entering its next phase of production deployments on a national scale with an elite team built from Palantir, Google, Meta, and Nvidia, and a growing footprint across government, insurance, health, and financial services.
Joining now means shaping both the company and a new category of applied AI. Every project here ships to production and is expected to create measurable customer value and impact.
You'll work alongside exceptional peers on some of the hardest problems in applied AI. It’s the kind of work you'll still be proud of in ten years from now.
Machine Learning Engineer, Applied AI
About the Role
So much of the work society depends on is still slower and harder than it should be. Permits take months. Claims sit unresolved. And AI hasn't changed that — because the bottleneck isn't the models. It's the institutional context AI needs to do the work: rules, history, relationships, and judgment scattered across people, documents, and legacy systems.
BrainCo exists to fix that. We build agent-native operating systems for the institutions society depends on, and our products are the first of their kind in the world — we were the first, anywhere, to fully automate construction permitting, and we're now doing the same across insurance and other industries. There is no playbook here, because no one has built this before.
As a Machine Learning Engineer on Applied AI, your work begins where the demo ends: getting a model to look impressive is the easy part; making it a production decision system an institution stakes its process on is the job. The problems come in every shape — custom vision model pipelines that check blueprints against building codes at 95%+ accuracy, agents that untangle policy stacks to reveal coverage gaps, systems that predict from clinical records whether a patient is on their care path — and you'll own them end-to-end, from ambiguous customer problem to the eval that catches a whole class of errors.
This is frontier ML applied where it's hardest and matters most. The problems are underspecified, the documents are brutal, the accuracy bar is institutional-grade — and the feedback loops are real, because our systems move real workflows forward every day.
Who We're Looking For
You understand how machine learning actually works — not just the tooling, but the philosophy underneath: what a loss function really optimizes, how generalization breaks under distribution shift, why evaluation is where systems quietly go wrong. And you live at the bleeding edge of modern AI, with hard-won instincts for squeezing the most out of LLMs and agentic systems — prompting, fine-tuning, tool use, and reasoning. That combination is the job: you know when a fine-tuned segmentation model beats a VLM, when a rule engine beats both, and how to compose all three into a system more accurate than any single model. You treat frontier models as components to be measured, pushed, and engineered — never as magic.
Most of all, you're energized by building things that have never existed, and comfortable when the problem, the data, and the definition of success all have to be invented at once.
The Problems You'll Work On
Composite AI systems and credit assignment. Our most demanding systems chain vision transformers, segmentation models, VLM reasoning, and rule engines. When the pipeline is wrong, which component failed? One of the most interesting open problems in applied ML.
Document understanding beyond the frontier. Blueprints, site plans, policy stacks, contracts, clinical records — dense, multimodal documents that break off-the-shelf models. You'll build models that actually read them.
Agents that learn from real work. Our deployments generate verified, ground-truth outcomes on every decision — reward signals most labs can only simulate. You'll help design the data, evals, and training loops to build and fine-tune agents on them.
Evaluation as a product discipline. When a regulator has to trust your system, evals are the product. You'll build eval suites and failure-mode taxonomies rigorous enough to earn institutional sign-off.
Institutional Intelligence that compounds. Every verified correction improves the system twice: the corrected fact percolates to every application, and the system that builds the intelligence learns to build it better. You'll work on both loops.
In This Role, You Will:
Turn ambiguity into shipped systems — from no problem statement, no labeled data, and no agreed definition of success, to well-posed ML problems and production deployments.
Own AI systems end-to-end. There is no handoff: the person who trains the model owns its behavior in production.
Work at the research frontier with production stakes, applying LLMs, RL fine-tuning, and agentic systems where the output is a decision an institution acts on.
Work directly with the institutions we serve — permit reviewers, underwriters, compliance officers — to understand how decisions actually get made and ensure your systems change how the work gets done.
Engineer for production reality, navigating accuracy, latency, cost, and reliability in environments far messier than any benchmark.
Raise the bar across the company through design reviews, our internal paper club, and the shared playbook for AI systems institutions can trust.
Senior Engineer Site Reliability
MatrixJV Co. Ltd.•Abu Dhabi, AE
5+ Years Exp
Posted: 15/9/2026
About the Company:
AIQ is an Abu Dhabi-based technology company that develops and deploys industrial artificial intelligence (AI) technologies at scale, focused on the energy sector. As a venture between Presight (G42) and ADNOC, AIQ has productized 15 AI-enabled solutions that support clients to perform better, protect teams and equipment, keep operations sustainable, and rapidly scale successes. The organization embraces an innovative and entrepreneurial spirit, looking to push boundaries and solve transformational challenges across industry. It welcomes professionals that share the desire to make meaningful and impactful contributions to the mission to deliver responsible AI to the heart of industrial processes. Always on the forefront of technology, AIQ provides its talent with an environment to thrive and excel. Working at AIQ includes participating in some of the most significant industrial transformation projects, interacting with massive data pools, utilizing sophisticated AI infrastructure that is powered by a NVIDIA GPU cloud computing platform, and access to abundant computing, storage, and network resources made available from across the G42 ecosystem.
Overview:
Job Title: Senior Engineer - Site Reliability
Job Location: Abu Dhabi, UAE
About AIQ
AIQ is a joint venture company in the United Arab Emirates, majority-owned by Presight (an ADX-listed G42 company) alongside ADNOC and G42, which focuses on developing artificial intelligence technologies. AIQ develops and commercializes AI products and applications for the oil and gas industry. It aims in providing end-to-end solutions by using its data, cloud and talents to develop AI solutions that seek to reduce costs and generate revenue for its clients.
AIQ embodies an innovative and entrepreneurial spirit that embraces challenges to push boundaries and seeks to welcome professionals to its team that share the desire to make meaningful and impactful contributions to its mission. Always on the cutting edge of technology, AIQ provides its talent all the opportunities to thrive and excel. Working at AIQ includes dealing with massive data sets, an AI infrastructure that is powered by the latest NVIDIA GPU cloud computing platform and access to limitless computing, storage and network resources.
Responsibilities:
Key Responsibilities
As a Senior SRE, you will enhance the reliability and performance of our platforms. You will lead key reliability projects, improve observability, and respond to complex production incidents.
Key Responsibilities:
• Maintain and evolve monitoring, alerting, and incident response systems.
• Proactively find and fix performance bottlenecks and failure points.
• Contribute to infrastructure automation and deployment pipelines.
• Drive SLO/SLI adoption in collaboration with engineering teams.
• Lead root cause analysis and build preventive solutions.
• Mentor junior engineers and help scale operational excellence.
• Analyze service performance, identify bottlenecks, and provide measurable improvement plans.
• Maintain the environment’s health by continuously monitoring technical and business metrics, configuring alerts for potential issues, and proactively addressing risks to prevent disruption.
• Deploy application updates with minimal disruption to services
• Identify, evaluate, and conduct proof-of-concepts for new technologies.
• Contribute to the knowledge base.
• Review and modify CI/CD principles and service maturity iteratively, striving for continuous improvement
• Comply with QHSE (Quality Health Safety and Environment), Business Continuity, Information Security, Privacy, Risk, Compliance Management, and Governance of Organizations policies, procedures, plans, and related risk assessments.
Qualifications:
Requirements:
Qualifications:
• Bachelor’s Degree in Business Analytics, Data Science, Computer Science, Engineering, or a related field.
• Master's Degree is preferred.
Experience:
• +5 years in a SRE/DevOps/Sysadmin/Platform Engineer role
• +5 years of experience in managing Kubernetes clusters.
• +5 years of experience in configuring and using monitoring/observability platforms
• Familiarity with at least one type of database
Skills
• Solid experience with containerized environments (Docker, Kubernetes).
• Hands-on with CI/CD pipelines and automation tools.
• Proficiency in scripting languages (Python, Bash).
• Strong grasp of observability tools (Prometheus, Grafana, ELK, Sentry).
• Good knowledge of cloud platforms (Huawei Cloud, Azure preferred)
Mandatory skills:
• Strong background in Linux/Unix Administration
• Solid hands-on experience deploying and operating Kubernetes or Openshift clusters
• Experience configuring and maintaining monitoring and observability solutions
• Ability to troubleshoot and resolve complex production issues efficiently, including performing root
• cause analysis and restoring services quickly during high-pressure incidents or critical outages
• Experience in backing up and restoring various systems
• Working together with project managers and solution architects while serving as subject matter
• Experts
• Implementing basic network security (e.g. configuring VPCs, firewalls/security groups, etc.)
• Understand the dependencies of various GPU cards, and upgrade container images as needed in
• order to ensure compatibility
• Deploy and operate products provided by third party providers
• Creating releases together with the development team and deploying release packages to all required environments
Bonus Skills:
• Good understanding of typical system architecture and interaction between its components
• Experience automating tasks using infrastructure-as-code tools, e.g. Ansible, Terraform
• Thorough understanding of a company's systems, including auxiliary components like caching
• systems (e.g., Redis, Memcached) and message queues (e.g., RabbitMQ, Kafka)
• Good understanding of databases, e.g. Postgres, Elasticsearch, Clickhouse
• Basic scripting
• Working knowledge of OAuth 2.0, OpenID/OpenID-Connect, SAML 2.0, Kerberos, LDAP
What working at AIQ offers:
Culture: Encouraging initiative, work in an environment that is fast-paced and varied, surrounded by talented peers from around the world, who are similarly attracted to applying their skills to solve transformational challenges.
Career: Join a team in which your contribution is recognized and rewarded, while you are supported to operate at your peak performance.
Rewards: An attractive renumeration package that includes healthcare, education support for dependents, leave benefits, and more.
ITS Engineer AI NLP
Parsons Corporation•Dubai, AE
7+ Years Exp
Posted: 15/9/2026
In a world of possibilities, pursue one with endless opportunities. Imagine Next!
At Parsons, you can imagine a career where you thrive, work with exceptional people, and be yourself. Guided by our leadership vision of valuing people, embracing agility, and fostering growth, we cultivate an innovative culture that empowers you to achieve your full potential. Unleash your talent and redefine what’s possible.
Job Description:
In a world of possibilities, pursue one with endless opportunities. Imagine Next!
At Parsons, you can imagine a career where you thrive, work with exceptional people, and be yourself. Guided by our leadership vision of valuing people, embracing agility, and fostering growth, we cultivate an innovative culture that empowers you to achieve your full potential. Unleash your talent and redefine what’s possible.
Parsons is looking for an amazingly talented ITS Engineer (AI & NLP) to join our team! In this role you will support the Resident Engineer in all technical, administrative, and operational aspects related to the development, integration, testing, validation, and deployment of Artificial Intelligence (AI), Machine Learning (ML), Natural Language Processing (NLP), Generative AI, and Decision Support applications within the Advanced Traffic Management System (ATMS). The role includes overseeing AI-enabled traffic management functions, predictive analytics, incident detection and classification, traffic forecasting, operator assistance tools, chatbot and virtual assistant capabilities, knowledge management solutions, and AI-driven business intelligence platforms. The position requires reviewing system designs, algorithms, data architectures, training methodologies, model validation processes, cybersecurity requirements, and integration with ATMS and other ITS systems to ensure compliance with project requirements and successful operational deployment.
What You'll Be Doing:
Support and supervise ITS design input, specification compliance, installation, testing, and integration across major infrastructure projects, ensuring all works meet contractual and technical requirements. Review and recommend approval of FSDDs, shop drawings, material submittals, method statements, and test plans, while coordinating the installation of ITS devices, fiber networks, power systems, and related civil foundations. Conduct routine site inspections to verify workmanship, safety, and compliance with approved designs. Oversee and witness FAT, SAT, prototype tests, OTDR fiber tests, communication and power tests, and reliability testing, preparing and closing punch lists prior to handover. Maintain accurate documentation of approvals, changes, correspondence, and photographic records, supporting the Resident Engineer with variations, claims, as‑built verification, O&M documentation, and final system commissioning.
Some of the principal responsibilities the AI & NLP Systems Engineer is required to perform are:
• Provide technical supervision of AI, ML, NLP, Large Language Model (LLM), predictive analytics, and decision-support system development and deployment activities.
• Review and recommend approval of Functional System Design Documents (FSDD), AI solution architectures, data models, AI governance plans, model development methodologies, and integration designs.
• Review and approve AI-related integration and testing plans, including model validation, performance testing, user acceptance testing, reliability testing, explainability assessments, and cybersecurity testing.
• Ensure the Contractor's compliance with the Contract Documents, approved AI development methodologies, data governance requirements, cybersecurity requirements, and applicable standards.
• Review and approve technical submittals related to AI platforms, NLP engines, knowledge management systems, LLM integrations, data analytics environments, and decision-support applications.
• Monitor the development and implementation of AI-enabled traffic management applications, including traffic prediction, incident detection, congestion analytics, operational recommendations, automated reporting, asset analytics, and performance monitoring solutions.
• Coordinate integration activities between AI systems, ATMS platforms, traffic databases, field ITS devices, video analytics systems, C-ITS platforms, enterprise systems, and external data sources.
• Review data quality, training datasets, model performance metrics, accuracy reports, bias assessments, and model improvement recommendations.
• Witness AI model demonstrations, prototype evaluations, proof-of-concept activities, validation exercises, and operational readiness testing.
• Support the Resident Engineer in evaluating Variation Orders, Change Orders, software enhancements, and system modifications related to AI and data analytics functions.
• Review issue logs, corrective action reports, model retraining plans, deployment roadmaps, and operational support procedures.
• Monitor AI solution deployment, release management, model version control, rollback procedures, and business continuity provisions.
• Attend all AI and NLP related meetings and prepare technical comments, recommendations, and meeting records.
• Review correspondence and prepare responses for signature.
• Maintain status logs for AI development, integration, testing, deployment, and issue resolution activities.
• Provide punch lists and final acceptance recommendations for AI-related deliverables.
• Review and approve as-built documentation, model documentation, data dictionaries, administrator manuals, operating procedures, training materials, and final handover documentation.
What Required Skills You'll Bring:
• Minimum 7 years of experience in software systems, intelligent transportation systems, data analytics, artificial intelligence, or digital transformation projects.
• Minimum 3 years of experience in AI, Machine Learning, Data Science, NLP, LLM, Advanced Analytics, or Decision Support System implementation projects.
• Bachelor's Degree in Computer Science, Artificial Intelligence, Data Science, Software Engineering, Computer Engineering, Information Technology, or related discipline from an accredited institution.
• Strong understanding of:
• Artificial Intelligence and Machine Learning technologies
• Natural Language Processing (NLP) and Large Language Models (LLM)
• Predictive analytics and forecasting systems
• Intelligent Decision Support Systems
• Data science methodologies and model lifecycle management
• Data architecture, data governance, and data quality management
• Enterprise software architecture and system integration
• API integrations, databases, and cloud-based environments
• MLOps, model training, deployment, monitoring, and version control
• AI explainability, model validation, and performance measurement
• Cybersecurity and secure AI implementation practices
• Traffic management operations and ITS business processes
• Excellent knowledge of AI integration within transportation management, smart city, mobility, or operational technology environments.
• Strong understanding of business intelligence platforms, data visualization, dashboard development, and reporting solutions.
• Experience with AI-enabled video analytics, traffic prediction, incident detection, knowledge management systems, or transportation analytics is highly desirable.
• Familiarity with cloud, virtualized, and hybrid computing environments.
• Experience in supervision of enterprise AI deployment and large-scale software integration projects.
• Knowledge of RTA ITS systems, ATMS platforms, and transportation operations is highly desirable.
• Strong coordination, stakeholder management, documentation, and technical review capabilities.
• Excellent written and verbal communication skills.
• Professional certifications in AI, Data Science, Cloud Technologies, Analytics, or Systems Engineering are considered an advantage.
• Arabic language capability is a plus.
Parsons equally employs representation at all job levels no matter the race, color, religion, sex (including pregnancy), national origin, age, disability or genetic information.
We truly invest and care about our employee’s wellbeing and provide endless growth opportunities as the sky is the limit, so aim for the stars! Imagine next and join the Parsons quest—APPLY TODAY!
Parsons is aware of fraudulent recruitment practices. To learn more about recruitment fraud and how to report it, please refer to https://www.parsons.com/fraudulent-recruitment/.
Site Reliability Engineer SRE
Bhft•AE
0+ Years Exp
Posted: 14/9/2026
We are looking for a Site Reliability Engineer who will be responsible for ensuring the reliable operation of our platform working with metrics to improve production process efficiency and participating in testing new product versions.Responsibilities:
Production Stability Management:
Ensure continuous compliance with external regulatory requirements and internal standards including risk security technology and trader needs.
Support and automate validation and monitoring processes for adherence to necessary standards.Incident Monitoring & Management:
Develop and improve monitoring and alerting systems to detect anomalies in key production metrics.
Implement rapid response mechanisms and efficient solutions to maintain strategy performance.Release & Change Management:
enforce standards for managing releases and changes to minimize deployment risks.
Implement strict acceptance testing for all releases.Process Management:
Develop and maintain Standard Operating Procedures (SOPs) for the team manage task queues and organize shift schedules to ensure continuous support and high availability of trading strategies.Integration Projects:
Lead initiatives to connect with new exchanges brokers and trading platforms ensuring smooth and secure service integration.Technical Performance Optimization:
Continuously improve system availability resilience (MTTR MTBF) and latency reduction while optimizing data exchange performance and order routing to maximize profitability.Qualifications :
Requirements:
Deep understanding of trading processes and market microstructure including colocation trading on native exchange protocols and algorithmic trading.Experience in monitoring alerting systems and incident management for highload environments.Knowledge of regulatory compliance and security standards.Proficiency in monitoring and incident management tools such as Grafana ClickHouse Prometheus Opsgenie Grafana OnCall PagerDuty etc.Experience developing and managing SOPs and KPIs for service teams.Experience managing integration projects with brokers and exchanges.Strong technical skill set including:
Linux systems administration and optimization.TCP/UDP multicast networking.FIXbased and native exchange protocolsColocation infrastructure setup and management.Python scripting for automation and monitoring.English proficiency at C1 level or higher.Remote Work :
YesEmployment Type :
Fulltime Key Skills Kubernetes,FMEA,Continuous Improvement,Elasticsearch,Go,Root cause Analysis,Maximo,CMMS,Maintenance,Mechanical Engineering,Manufacturing,Troubleshooting Experience:
years Vacancy:
1
Service Delivery Manager
Dxc Technology•Abu Dhabi, AE
8+ Years Exp
Posted: 14/9/2026
DXC Technology | United Arab Emirates | On-site
DXC Technology is seeking an experienced Service Delivery Manager to lead and govern the delivery of complex, business-critical IT services within a large-scale technology environment in the UAE.
The Service Delivery Manager will be responsible for ensuring consistent service performance, operational excellence, customer satisfaction, SLA achievement, effective governance, and continuous improvement across multiple technology services, teams, vendors, and partners.
This role requires a strong combination of IT service management, operational leadership, stakeholder management, service governance, major incident oversight, and multi-vendor coordination.
Key Responsibilities
Service Delivery & Operational Management
• Own end-to-end service delivery across assigned IT services and technology domains.
• Ensure services are delivered in accordance with agreed SLAs, KPIs, OLAs, contractual commitments, and quality standards.
• Monitor service performance, availability, reliability, capacity, and operational health.
• Coordinate service delivery activities across internal teams, technical towers, vendors, and partners.
• Identify service delivery gaps and drive corrective and preventive actions.
• Maintain operational readiness and service continuity across the environment.
• Ensure effective transition of new and changed services into BAU operations.
IT Service Management
• Drive effective implementation and operation of ITIL-based service management processes.
• Provide governance across Incident, Major Incident, Problem, Change, Request, Configuration, Availability, Capacity, and Service Level Management.
• Ensure incidents and service requests are managed within agreed service levels.
• Drive root-cause analysis and permanent resolution of recurring service issues.
• Ensure effective change governance while minimizing operational risk and service disruption.
• Promote consistent service management standards across all delivery teams.
Major Incident & Problem Management
• Provide leadership and oversight during critical and major service incidents.
• Ensure appropriate technical teams, vendors, and stakeholders are mobilized for rapid service restoration.
• Maintain clear communication and escalation throughout major incidents.
• Ensure comprehensive post-incident reviews and root-cause analyses are completed.
• Track corrective and preventive actions through to closure.
• Identify recurring operational trends and drive permanent remediation.
Service Governance & Performance
• Establish and maintain effective service governance and reporting frameworks.
• Conduct regular service reviews with stakeholders, technical teams, vendors, and partners.
• Monitor and report performance against SLAs, KPIs, service quality measures, risks, and improvement plans.
• Produce clear management and executive-level service reports.
• Maintain service risks, issues, actions, and improvement plans.
• Ensure timely escalation and resolution of service performance concerns.
Stakeholder Management
• Act as a key service delivery interface between DXC Technology, stakeholders, delivery teams, and technology partners.
• Build trusted relationships with senior stakeholders and service owners.
• Manage expectations and ensure clear, transparent, and timely communication.
• Translate technical service issues into clear business impact and resolution plans.
• Coordinate across multiple technical and business teams to achieve service objectives.
• Maintain a strong focus on customer experience and satisfaction.
Vendor & Partner Management
• Manage service delivery across multiple vendors, partners, and technology providers.
• Monitor vendor performance against agreed service levels and contractual commitments.
• Coordinate resolution of cross-vendor incidents, problems, and operational dependencies.
• Drive accountability for service performance and improvement actions.
• Ensure effective collaboration across the wider service delivery ecosystem.
Service Improvement
• Establish and maintain Continual Service Improvement (CSI) plans.
• Analyze service performance data, trends, recurring incidents, and customer feedback to identify improvement opportunities.
• Drive initiatives to improve service quality, stability, availability, efficiency, and customer experience.
• Identify opportunities for automation, process optimization, and operational efficiency.
• Track improvement initiatives through to measurable outcomes.
Service Continuity & Resilience
• Support service continuity, availability, resilience, and disaster recovery requirements.
• Ensure operational teams maintain appropriate procedures, runbooks, escalation paths, and recovery processes.
• Participate in service continuity and disaster recovery exercises where required.
• Ensure operational risks and service vulnerabilities are identified and appropriately managed.
Required Experience
• Bachelor's degree in Information Technology, Computer Science, Engineering, Business, or a related discipline.
• 8+ years of IT service delivery / IT operations experience, including significant experience in a Service Delivery Manager or equivalent role.
• Proven experience managing complex enterprise IT services across multiple technology domains.
• Strong knowledge of ITIL and IT Service Management (ITSM) principles and processes.
• Demonstrated experience managing SLAs, KPIs, service performance, governance, and customer satisfaction.
• Strong experience in Incident, Major Incident, Problem, Change, and Service Level Management.
• Experience managing services within complex multi-vendor environments.
• Proven experience managing senior stakeholders and service review/governance meetings.
• Strong understanding of enterprise IT operations, infrastructure, applications, cloud, network, security, and service support environments.
• Experience with ITSM platforms such as ServiceNow or equivalent.
• Excellent leadership, communication, stakeholder management, analytical, and problem-solving capabilities.
Preferred Experience
• Previous experience managing large-scale technology services in the UAE or wider GCC region.
• Experience supporting complex, business-critical or high-availability technology environments.
• Experience working within large-scale, highly governed enterprise environments.
• Experience managing service transition from project delivery into BAU operations.
• Understanding of service continuity, disaster recovery, availability, capacity, and operational resilience.
• Experience managing multiple technology vendors and service providers.
• ITIL 4 Foundation, ITIL Managing Professional, or equivalent ITSM certification is highly desirable.
• Relevant ServiceNow, cloud, infrastructure, or service management certifications are advantageous.
What We Are Looking For
We are looking for a hands-on Service Delivery Manager who can maintain strong operational control while building trusted relationships with stakeholders.
The successful candidate will understand how complex technology services operate end-to-end and will be able to coordinate multiple technical teams and vendors to maintain service quality, stability, availability, and performance.
The ideal candidate will combine strong ITSM knowledge, operational leadership, stakeholder management, service governance, and a continuous-improvement mindset.
Join DXC Technology and play a key role in delivering reliable, high-quality technology services in the UAE.
Senior Site Reliability Engineer
Aift•Sharjah, AE
5+ Years Exp
Posted: 12/9/2026
Job Overview
We are looking for a hands-on infrastructure engineer to own the deployment, migration and troubleshooting of on-premise Kubernetes environments for enterprise and government clients, including airgapped, high-security data center environments where remote access is not possible.
This is a client-facing, on-site role: you will be the technical authority in the room, responsible for executing complex infrastructure changes correctly the first time, diagnosing failures independently under pressure and communicating clearly with client stakeholders throughout.
This role carries real ownership; you will be expected to understand the systems deeply enough to make sound judgment calls when things don't go to plan, without waiting on remote support.
In this role, you will play a vital part in supporting our Cybersecurity business, Vulcan. Vulcan is a cybersecurity solution for GenAI, providing red and blue team services to ensure compliance and security.
Learn more about us
Vulcan product: https://vulcanlab.ai/
Vulcan LinkedIn: https://www.linkedin.com/company/vulcanlab-ai/
AIFT group: https://aift.io/
Responsibilities
Plan and execute on-prem Kubernetes cluster deployments, upgrades and infrastructure migrations (including IP re-addressing, certificate rotation and cluster reconfiguration) in production and airgapped environments
Diagnose and resolve failures independently on-site
Own the full infrastructure stack end-to-end: Kubernetes control plane and data plane, PostgreSQL (primary/replica replication), distributed storage (e.g. SeaweedFS/Ceph/similar), private container registries and centralized logging (ELK or equivalent)
Validate deployment tooling (scripts, installers, automation) thoroughly in lab/staging environments before any client-facing execution
Represent the technical work directly to client stakeholders on-site: explain status, failures and remediation plans clearly
Travel to client data centers (including airgapped/restricted-access sites) as , sometimes on short notice, for deployment and go-live support
Write clear, structured runbooks, decision trees and incident reports that others (including less experienced engineers) can follow under pressure
Escalate risks proactively to internal leadership, not just after something has gone wrong
Requirements
Technical:
5-6 years of hands-on experience with Kubernetes in production, including at least one on-premise (not purely cloud-managed) deployment
Solid understanding of etcd internals. Quorum, peer membership, failure recovery, not just kubectl-level familiarity
Experience with kubeadm-based cluster bootstrapping and certificate management (SANs, CA rotation, renewal)
Working knowledge of PostgreSQL replication, Linux networking fundamentals (DNS, NTP, firewalls) and container registries (Docker Distribution or similar)
Comfortable working entirely from the Linux command line, writing and debugging bash scripts and reading unfamiliar automation tooling under time pressure
Experience with at least one distributed storage system (SeaweedFS, Ceph, MinIO or similar) is a strong plus
GPU-enabled Kubernetes nodes (NVIDIA device plugin, container toolkit) experience is a plus, not required
Working style:
Demonstrated ability to work independently in high-pressure, high-stakes environments without live support
Strong incident communication. Can explain technical failures to non-technical stakeholders factually and calmly, without over-promising or minimizing
A track record of validating changes in test environments before touching production and the judgment to insist on this even under deadline pressure
Comfortable with travel, including to secure/restricted facilities where personal devices, internet access or remote assistance may not be available
Nice to have:
Prior consulting, systems integration or professional services experience, ideally on enterprise or government accounts
Experience specifically in the GCC/Middle East region, or with government-sector clients
Security background (the ability to reason about access controls, credential handling and airgapped operational discipline is valuable given the environments involved)
Interview Process
HR phone interview: 1 hour
Online interview: 1.5 - 2 hours, meet with hiring manager
Online interview: 1 hour, meet with hiring team
Why Join Us?
Innovative Environment:
Be part of a company at the forefront of technology to provide security in GenAI, with opportunities to work on groundbreaking projects.
Growth Opportunities:
Take your career to new heights with our career development programs and growth-focused culture.
Dynamic Team:
Join a multi-cultural and dynamic team of dedicated professionals who inspire and support each other.
Compensation
:
Competitive salary and benefits package, commensurate with experience and performance.
Senior Site Reliability Engineer
Aift•Fujairah, AE
5+ Years Exp
Posted: 12/9/2026
Job Overview
We are looking for a hands-on infrastructure engineer to own the deployment, migration and troubleshooting of on-premise Kubernetes environments for enterprise and government clients, including airgapped, high-security data center environments where remote access is not possible.
This is a client-facing, on-site role: you will be the technical authority in the room, responsible for executing complex infrastructure changes correctly the first time, diagnosing failures independently under pressure and communicating clearly with client stakeholders throughout.
This role carries real ownership; you will be expected to understand the systems deeply enough to make sound judgment calls when things don't go to plan, without waiting on remote support.
In this role, you will play a vital part in supporting our Cybersecurity business, Vulcan. Vulcan is a cybersecurity solution for GenAI, providing red and blue team services to ensure compliance and security.
Learn more about us
Vulcan product: https://vulcanlab.ai/
Vulcan LinkedIn: https://www.linkedin.com/company/vulcanlab-ai/
AIFT group: https://aift.io/
Responsibilities
Plan and execute on-prem Kubernetes cluster deployments, upgrades and infrastructure migrations (including IP re-addressing, certificate rotation and cluster reconfiguration) in production and airgapped environments
Diagnose and resolve failures independently on-site
Own the full infrastructure stack end-to-end: Kubernetes control plane and data plane, PostgreSQL (primary/replica replication), distributed storage (e.g. SeaweedFS/Ceph/similar), private container registries and centralized logging (ELK or equivalent)
Validate deployment tooling (scripts, installers, automation) thoroughly in lab/staging environments before any client-facing execution
Represent the technical work directly to client stakeholders on-site: explain status, failures and remediation plans clearly
Travel to client data centers (including airgapped/restricted-access sites) as , sometimes on short notice, for deployment and go-live support
Write clear, structured runbooks, decision trees and incident reports that others (including less experienced engineers) can follow under pressure
Escalate risks proactively to internal leadership, not just after something has gone wrong
Requirements
Technical:
5-6 years of hands-on experience with Kubernetes in production, including at least one on-premise (not purely cloud-managed) deployment
Solid understanding of etcd internals. Quorum, peer membership, failure recovery, not just kubectl-level familiarity
Experience with kubeadm-based cluster bootstrapping and certificate management (SANs, CA rotation, renewal)
Working knowledge of PostgreSQL replication, Linux networking fundamentals (DNS, NTP, firewalls) and container registries (Docker Distribution or similar)
Comfortable working entirely from the Linux command line, writing and debugging bash scripts and reading unfamiliar automation tooling under time pressure
Experience with at least one distributed storage system (SeaweedFS, Ceph, MinIO or similar) is a strong plus
GPU-enabled Kubernetes nodes (NVIDIA device plugin, container toolkit) experience is a plus, not required
Working style:
Demonstrated ability to work independently in high-pressure, high-stakes environments without live support
Strong incident communication. Can explain technical failures to non-technical stakeholders factually and calmly, without over-promising or minimizing
A track record of validating changes in test environments before touching production and the judgment to insist on this even under deadline pressure
Comfortable with travel, including to secure/restricted facilities where personal devices, internet access or remote assistance may not be available
Nice to have:
Prior consulting, systems integration or professional services experience, ideally on enterprise or government accounts
Experience specifically in the GCC/Middle East region, or with government-sector clients
Security background (the ability to reason about access controls, credential handling and airgapped operational discipline is valuable given the environments involved)
Interview Process
HR phone interview: 1 hour
Online interview: 1.5 - 2 hours, meet with hiring manager
Online interview: 1 hour, meet with hiring team
Why Join Us?
Innovative Environment:
Be part of a company at the forefront of technology to provide security in GenAI, with opportunities to work on groundbreaking projects.
Growth Opportunities:
Take your career to new heights with our career development programs and growth-focused culture.
Dynamic Team:
Join a multi-cultural and dynamic team of dedicated professionals who inspire and support each other.
Compensation
:
Competitive salary and benefits package, commensurate with experience and performance.
Senior DevOps Engineer
GSSTech Group•Dubai, AE
0+ Years Exp
Posted: 8/9/2026
About the Role
We are seeking an experienced and highly skilled Senior DevOps / SRE Engineer to design, implement, automate, and manage modern DevOps and Site Reliability Engineering practices across enterprise technology environments. The ideal candidate will possess strong hands-on experience in CI/CD, GitHub Actions, Jenkins, container technologies, Infrastructure as Code (IaC), DevSecOps, monitoring, logging, automation, and cloud-native technologies. This role requires collaboration across development, infrastructure, security, and operations teams to enhance deployment efficiency, system reliability, security, scalability, and operational excellence. The successful candidate will demonstrate ownership of projects from inception through delivery and ongoing operational support while continuously identifying opportunities to automate and improve existing processes.
Key Responsibilities
DevOps & CI/CD
• Design, implement, maintain, and optimize CI/CD pipelines for enterprise applications and services.
• Develop and manage deployment pipelines using Jenkins and GitHub Actions.
• Build and maintain Jenkins pipelines using:
• Jenkins Shared Libraries
• Declarative and Scripted Pipelines
• Groovy scripting
• Python and Bash automation
• Implement automated build, test, security scanning, packaging, and deployment processes.
• Develop reusable CI/CD components and standards to improve consistency across development teams.
• Support automated deployments across containerized and cloud-native environments.
• Identify and eliminate manual deployment activities through automation.
• Establish best practices around source control, branching, release management, artifact management, and deployment automation.
DevOps Tooling
• Install, configure, administer, upgrade, integrate, and troubleshoot enterprise DevOps tooling, including:
• Jenkins
• GitHub / GitHub Actions
• Nexus / Nexus IQ
• SonarQube
• Checkmarx
• Sysdig
• Cosign
• Argo CD
• JIRA
• Confluence
• Manage integrations between DevOps tools to establish an efficient and secure software delivery lifecycle.
• Monitor tool availability, performance, capacity, and security.
• Troubleshoot issues related to CI/CD tools and their integrations.
• Establish standards for tool configuration, access management, security, and governance.
Containerization & Deployment
• Implement and manage CI/CD pipelines for applications deployed on container technologies.
• Support containerized application build, packaging, deployment, and lifecycle management.
• Work closely with development and infrastructure teams to improve container-based application delivery.
• Implement automated deployment strategies using modern DevOps and GitOps practices.
• Support Argo CD-based GitOps workflows and ensure reliable application deployments.
• Troubleshoot container, deployment, configuration, and runtime-related issues.
Infrastructure Automation & Configuration Management
• Automate infrastructure provisioning, deployment, and configuration management activities.
• Develop reusable automation to create consistent, scalable, and repeatable environments.
• Apply Infrastructure as Code principles to infrastructure and application configuration.
• Reduce operational dependency on manual activities through automation.
• Implement configuration management standards and ensure consistency across environments.
• Support infrastructure changes through controlled, automated, and auditable processes.
DevSecOps & Security
• Integrate security controls into CI/CD pipelines following DevSecOps principles.
• Implement and maintain automated security and quality gates using tools such as:
• SonarQube
• Checkmarx
• Nexus IQ
• Sysdig
• Cosign
• Support vulnerability scanning, code quality analysis, dependency analysis, container security, and artifact verification.
• Implement secure software supply-chain practices, including artifact signing and verification.
• Work with security teams to address vulnerabilities and improve application and infrastructure security.
• Ensure DevOps processes comply with organizational security and governance standards.
Monitoring, Logging & Observability
• Implement and maintain enterprise logging, monitoring, and observability solutions.
• Work with technologies such as:
• ELK Stack
• Prometheus
• Grafana
• Develop dashboards, alerts, and monitoring mechanisms for infrastructure and application environments.
• Monitor system health, application performance, availability, and reliability.
• Analyze logs and metrics to identify performance issues and operational risks.
• Establish proactive monitoring and alerting to reduce incidents and improve service reliability.
• Support root-cause analysis of production incidents using monitoring and observability data.
Incident Management & Troubleshooting
• Troubleshoot and resolve complex infrastructure, application, deployment, and CI/CD issues.
• Work closely with development, infrastructure, security, and operations teams to resolve production and non-production issues.
• Participate in incident management, problem management, and root-cause analysis.
• Identify recurring issues and implement permanent solutions rather than relying on manual workarounds.
• Support production deployments and provide operational assistance when required.
• Contribute to continuous improvement initiatives based on incident trends and operational feedback.
SRE & Operational Excellence
• Apply Site Reliability Engineering (SRE) principles to improve system availability, scalability, performance, and resilience.
• Implement automation to reduce operational toil.
• Define and improve operational processes, reliability practices, and service standards.
• Support capacity planning, performance optimization, availability, and resilience initiatives.
• Contribute to reliability engineering practices such as monitoring, alerting, incident response, and continuous improvement.
• Balance project delivery responsibilities with operational support requirements.
DevOps & Application Lifecycle
• Work on projects from inception, design, implementation, testing, deployment, and transition to operations.
• Collaborate with application development teams to integrate DevOps practices into the software development lifecycle.
• Provide technical guidance on build, deployment, configuration, automation, and operational requirements.
• Promote automation-first approaches across development and operations teams.
• Support continuous improvement of engineering processes and delivery methodologies.
Required Technical Skills
Mandatory Skills
• Strong hands-on experience with DevOps methodologies and practices.
• Strong practical experience with GitHub Actions.
• Strong hands-on experience with Jenkins and CI/CD pipelines.
• Good knowledge of Jenkins Shared Libraries, Groovy, Python, and/or Bash scripting.
• Experience managing and administering enterprise DevOps tooling.
• Strong understanding of container technologies and container-based deployments.
• Experience with GitOps and Argo CD.
• Strong understanding of Infrastructure as Code (IaC) principles.
• Strong knowledge of DevSecOps and security integration within CI/CD pipelines.
• Experience with ELK, Prometheus, and Grafana or equivalent monitoring and logging technologies.
• Strong infrastructure and security fundamentals.
• Strong troubleshooting and problem-solving capabilities.
DevOps Tools
Hands-on experience with several of the following:
• Jenkins
• GitHub
• GitHub Actions
• Nexus
• Nexus IQ
• SonarQube
• Checkmarx
• Sysdig
• Cosign
• Argo CD
• JIRA
• Confluence
DevOps / SRE Knowledge
The candidate should have a comprehensive understanding of:
• DevOps principles and methodologies
• SRE principles and practices
• CI/CD and continuous delivery
• Infrastructure as Code
• GitOps
• DevSecOps
• Containerization
• Automation and configuration management
• 12-Factor Application principles
• Monitoring and observability
• Logging and incident management
• Application lifecycle management
• Infrastructure and application security
• Reliability, scalability, and availability engineering
Soft Skills & Behavioral Competencies
• Strong communication skills with the ability to communicate effectively with technical and non-technical stakeholders.
• Ability to establish transparent and professional relationships with internal teams, clients, vendors, and key stakeholders.
• Strong ownership and accountability for assigned projects and services.
• Ability to work independently while collaborating effectively within cross-functional teams.
• Strong analytical and pragmatic approach to problem-solving.
• Ability to work effectively under pressure and manage multiple priorities.
• Strong stakeholder and client management skills.
• Ability to manage client expectations while maintaining delivery commitments.
• Strong focus on operational excellence and continuous improvement.
• Willingness to learn and adopt emerging technologies and industry best practices.
• Ability to work across both project delivery and operational support functions.
• Ability to adapt to changing business and technology requirements.
• Demonstrate professional integrity, accountability, and a collaborative approach.
• Ability to contribute positively to team culture and promote organizational values.
Key Success Factors
The successful candidate will be expected to:
• Increase deployment automation and reduce manual intervention.
• Improve CI/CD pipeline reliability and efficiency.
• Strengthen security throughout the software delivery lifecycle.
• Improve infrastructure and application reliability.
• Reduce recurring production incidents through automation and root-cause analysis.
• Improve monitoring, logging, and observability.
• Establish scalable and reusable DevOps practices.
• Support faster and more reliable software delivery.
• Build strong relationships with development, infrastructure, security, operations, and business stakeholders.
• Continuously identify opportunities for process, tooling, and technology improvements.
Vice President Site Reliability Engineering
Remotedxb•Dubai, AE
6+ Years Exp
Posted: 6/9/2026
Responsibilities
• Oversee a specialized SRE team focused on the design, deployment, and maintenance of automation toolsets
• Establish and enforce standards for Infrastructure as Code (IaC) using Terraform
• Lead strategy for automated configuration and state management using Ansible and Packer
• Manage monitoring and health of automation platforms using SLIs/SLOs
• Drive automated lifecycle management for physical and virtual assets
• Lead development of custom scripts and internal providers using Python, Go, PowerShell, or Bash
• Collaborate with the Datacenter team to facilitate workflows and system needs
• Analyze system behavior and resource utilization to optimize automated deployments
• Provide technical guidance and career mentorship to SREs
Requirements
• 6-10 years of experience in Infrastructure, SRE, or DevOps focused on automation at scale
• Deep proficiency with Terraform and Ansible
• Hands‑on experience with image creation using Packer, Ansible, or SCCM
• Experience managing VMware (vSphere/vCenter) and cloud providers like Azure and AWS
• High‑level scripting skills in Python, Go, PowerShell, and Bash
• Experience with observability tools such as Splunk, ELK, Prometheus, or Grafana
• Understanding of network topology and experience with Juniper or Palo Alto
• Mastery of Git and CI/CD platforms like Jenkins, GitLab CI, or GitHub Actions
• Proficiency in managing both Windows Server and Linux
Preferred Qualifications
• Previous experience in team leadership or management
• Experience with IAM platforms like Entra ID, Active Directory, or Okta
• Experience with block or object storage (HP Alletra, EMC, DDN, S3, Azure Blob)
• Experience with storage backup and DR management using Commvault or Veeam
About the Company
Galaxy is a global leader in digital assets and data center infrastructure, delivering solutions that accelerate progress in finance and artificial intelligence.
#J-18808-Ljbffr