Senior AI Platform Engineer (Cloud) - Sandton
- Location
- Johannesburg, Gauteng
- Closing date
First listed . Last checked at source .
In brief
**Overview** Absa Group’s Chief Data Analytics and Applied AI Office (CDAIO) is seeking a Senior AI Platform Engineer (Cloud) to be based in Sandton. This role is responsible for designing, building, operating, and continuously optimizing the multi-cloud AI infrastructure that powers the bank's enterprise AI capability. This infrastructure enables the CDAIO to fulfill its mandate as the steward of the bank’s AI capabilities through end-to-end delivery of the AI platform enablement, governance, and acceptable use in service of the bank’s strategic and commercial objectives. The position serves as the engineering backbone for a platform supporting various live AI projects across four business units (CIB, PPB, BB, AR) and ten countries. It requires deep technical mastery in cloud AI infrastructure, AI FinOps, zero-trust security architecture, agentic AI infrastructure, and platform observability, combined with the commercial fluency to govern AI compute costs at enterprise scale and communicate trade-offs to senior business and finance stakeholders. The role involves applying critical thinking, design thinking, and problem-solving skills in an agile team environment to solve complex platform engineering challenges, delivering high-quality, cost-optimal solutions in full compliance with Absa's Enterprise-Wide Risk Management Framework, Group Architecture standards, and AI Responsible Use Policy. The successful candidate will hold full accountability for building high-performing, scalable, enterprise-grade Platform services and for developing capability in others. **What you will do** * Lead the design, deployment, and continuous optimisation of Absa's multi-cloud AI platform stack including AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face Model Hub, and on-demand GPU clusters. * Architect scalable, resilient, and reusable platform components such as AI Gateway configuration, model serving infrastructure, vector database deployments, and data pipeline integration. * Define and maintain infrastructure-as-code (IaC) standards, such as using Terraform or Pulumi, for repeatable, auditable multi-cloud AI deployments. * Lead the design and operation of agentic AI infrastructure, including orchestration runtime environments (e.g., Microsoft Foundry Agent Service, AWS Bedrock Agents), tool-calling schemas, agent memory and state management patterns, and multi-agent communication protocols. * Develop and enforce cloud-agnostic model serving patterns to reduce platform lock-in and ensure workload portability. * Identify and select appropriate internal and external technologies to deliver AI platform services and continuously improve platform engineering practices. * Take full accountability for end-to-end platform quality, completeness, and user experience. * Positively contribute to the design and evolution of Group Architecture, infrastructure standards, and AI platform governance frameworks. * Own the AI compute cost model for the CDAIO, including chargeback and showback frameworks for various AI service consumptions. * Design and maintain FinOps dashboards and cost attribution reports using tools like AWS Cost Explorer, Databricks System Tables cost analytics, and Azure OpenAI utilisation tooling. * Evaluate and manage provisioned throughput versus on-demand consumption trade-offs for production AI workloads, presenting optimisation recommendations. * Identify and execute AI compute cost optimisation opportunities, such as workload scheduling, spot instance strategies for training workloads, model distillation, and right-sizing of GPU clusters. * Create business cases and solution specifications for AI platform investments and governance processes. * Collaborate with the FinOps capability within the CDAIO COO to align AI platform costs to agreed budget envelopes and ensure proactive detection and escalation of spend anomalies. * Define, implement, and own AI-specific SLAs and OLAs covering inference latency, platform availability, token throughput, API gateway response times, and model serving reliability. * Implement and maintain AI platform observability tooling (e.g., Prometheus, Grafana, Datadog, Databricks Lakehouse Monitoring) for real-time visibility of platform health, model drift alerts, and capacity utilisation. * Design and operate incident management processes for AI platform failures, including on-call runbooks, escalation paths, post-incident reviews, and root-cause remediation. * Lead service improvement initiatives, translating performance data into platform enhancement programmes and continuously reducing mean time to recovery (MTTR). * Own the release and change management process for AI platform components, including change governance, cutover management, and operational readiness sign-off. * Use production performance monitoring and customer data to inform technical design and implementation decisions. * Design and implement zero-trust security architecture for AI platform APIs and services (e.g., OAuth 2.0 / OIDC integration, JWT/JWE/JWS token management, RBAC, ABAC). * Implement prompt injection prevention, output filtering, and data exfiltration controls at the AI Gateway layer. * Design and enforce data residency and sovereignty controls for AI platform deployments across Absa's operating countries. * Conduct and maintain AI-specific threat models in collaboration with the Chief Information Security Office. * Apply and maintain all Group risk, governance, compliance, and regulatory standards and frameworks; hold accountability for all risk associated with AI platform engineering decision-making. * Update, develop, and maintain all platform documentation in accordance with organisational technical standards and risk and governance frameworks. * Lead and develop a team of AI Platform Engineers, establishing clear performance objectives, providing regular coaching and feedback. * Cascade platform direction across the team, ensuring alignment on platform strategy, performance objectives, and delivery priorities. * Leverage coaching techniques across all squad-related activity to drive higher-quality design and deployment of AI platform services. * Maintain comprehensive technical documentation, architectural decision records (ADRs), and operational runbooks for all platform components. * Conduct peer reviews, testing, and problem-solving within and across the broader CDAIO engineering community; identify and develop needed skills in self and others. * Support the AI Embedment and Training capability in developing platform onboarding materials and self-service guides. * Proactively lead agile practices, remove barriers to success, and ensure seamless delivery in a continuously changing environment. **Requirements (from the original advert)** **Education/ Qualification:** * Postgraduate degree in a quantitative discipline such as Computer Science, Data Science, Mathematics, Statistics, Engineering, or equivalent (Masters-essential or PhD-advantageous). * Bachelor's Degree: Information Technology. **Certification in:** * Cloud: AWS Solutions Architect Professional, AWS Machine Learning Specialty, or Microsoft Azure AI Engineer Associate. * FinOps: FinOps Foundation Certified Practitioner (FOCP) or equivalent AI cost governance credential. * Security Certification: Certified Cloud Security Professional (CCSP) or AWS Security Specialty. * IaC Certification: HashiCorp Terraform Associate or equivalent infrastructure-as-code credential. **Work Experience:** * 5-8 years of progressive leadership experience in Cloud AI Platform Engineering, with production experience managing multi-cloud AI platform stacks across at least two of: AWS Bedrock/SageMaker, Databricks AI, Microsoft Azure AI Foundry, or Hugging Face enterprise deployments. **Minimum 2-3 years experience in the following:** * AI FinOps and Cost Governance: Demonstrated ownership of AI compute cost models and FinOps reporting in a multi-BU or multi-cloud environment, with evidence of cost optimisation outcomes. * AI Security Architecture: Designing and implementing zero-trust AI security (OAuth/OIDC, JWT, prompt injection controls, data residency compliance) in a regulated environment. * Agentic AI Infrastructure: Production design of agent orchestration infrastructure such as LangGraph, AutoGen, Foundry Agent Service, Bedrock Agents, tool-calling APIs, and agent state management. * Platform Observability: Operating AI-specific observability tooling for inference latency, drift alerting, and capacity management such as Prometheus, Grafana, Datadog, or Lakehouse Monitoring. * Infrastructure-as-Code: Terraform, Pulumi, or equivalent for multi-cloud, multi-region AI infrastructure deployments; CI/CD pipeline design for platform components. * Regulated Industry: AI platform engineering in financial services or a similarly regulated sector with model risk governance and change management obligations. **Advantageous:** * People leadership: Leading or mentoring a team of platform or infrastructure engineers in an agile delivery environment. * Pan-African Deployments: Delivering AI platform services across multiple African jurisdictions with awareness of data localisation and cross-border data transfer requirements. **Knowledge and Skills:** * Multi-Cloud AI Platform Architecture: Expert design and operation of AWS Bedrock, Databricks AI, Azure AI Foundry, and Hugging Face in enterprise production environments across multiple business units and geographies. * Agentic AI Infrastructure: Practical production knowledge of agent orchestration frameworks (LangGraph, AutoGen, Foundry Agent Service, Bedrock Agents), tool-calling API design, agent memory architecture, and multi-agent coordination patterns. * AI FinOps and Cost Management: Chargeback and showback model design; DBU and token cost attribution; provisioned throughput versus on-demand optimisation; GPU cluster cost management; spend anomaly detection and FinOps dashboarding. * AI Security and Zero Trust: OAuth 2.0, OIDC, JWT/JWE/JWS; RBAC and ABAC for AI workloads; prompt injection prevention; data exfiltration controls at the Gateway layer; AI threat modelling and data residency compliance. * Infrastructure-as-Code: Terraform, Pulumi, or AWS CDK for multi-cloud AI infrastructure; CI/CD pipeline design for platform components; container orchestration using Docker, Kubernetes, and Helm. * Platform Observability: Prometheus, Grafana, Datadog, OpenTelemetry, and Databricks Lakehouse Monitoring; custom metric design for AI workload health including inference latency, token throughput, and model drift. * Cloud-Agnostic Model Serving: ONNX, BentoML, Triton Inference Server; containerised model deployment patterns for portability across AWS, Azure, and Databricks environments. * MLOps Tooling: Working knowledge of MLflow, Kubeflow, Airflow, and CI/CD for ML, sufficient to collaborate effectively with AI Solution Engineers on model deployment and lifecycle management. * GPU and HPC Architecture: On-demand GPU cluster management; spot instance strategies; high-performance compute cost optimisation for large-scale model training and fine-tuning workloads. * Enterprise Risk and Governance: Absa Enterprise Wide Risk Management Framework; Group Architecture standards; AI Responsible Use Policy; POPIA; country-specific data localisation requirements across Absa's ten operating countries. * Agile Delivery: Sprint planning, backlog management, and continuous delivery practices in a self-directed squad environment; experience removing delivery barriers in a fast-moving, multi-stakeholder context. **Who should apply** The ideal candidate is a technically exceptional and commercially grounded AI Platform Engineer with deep technical mastery in cloud AI infrastructure, AI FinOps, zero-trust security architecture, agentic AI infrastructure, and platform observability. They must possess commercial fluency to govern AI compute costs at an enterprise scale and effectively communicate trade-offs to senior business and finance stakeholders. The candidate should demonstrate critical thinking, design thinking, and problem-solving skills within an agile team environment, capable of solving complex platform engineering challenges and delivering high-quality, cost-optimal solutions. This role requires an individual who can take full accountability for building high-performing, scalable, enterprise-grade Platform services and for developing capabilities in others. Applicants must hold a postgraduate degree in a quantitative discipline (Masters-essential or PhD-advantageous) and a Bachelor's Degree in Information Technology. Required certifications include AWS Solutions Architect Professional, AWS Machine Learning Specialty, or Microsoft Azure AI Engineer Associate; FinOps Foundation Certified Practitioner (FOCP) or equivalent; Certified Cloud Security Professional (CCSP) or AWS Security Specialty; and HashiCorp Terraform Associate or equivalent. Candidates should have 5-8 years of progressive leadership experience in Cloud AI Platform Engineering, with specific production experience in multi-cloud AI platform stacks. Additionally, 2-3 years of experience is required in AI FinOps and Cost Governance, AI Security Architecture, Agentic AI Infrastructure, Platform Observability, Infrastructure-as-Code, and within a regulated industry. Proficiency in a comprehensive list of knowledge and skills including multi-cloud AI platform architecture, agentic AI infrastructure, AI FinOps, AI Security, IaC, platform observability, cloud-agnostic model serving, MLOps tooling, GPU/HPC architecture, enterprise risk, and agile delivery is essential. Experience in people leadership and Pan-African deployments is advantageous. **Deadline** October 16, 2026 **Reference** Original posting: https://www.myjobmag.co.za/job/senior-ai-platform-engineer-cloud-sandton-absa-group-limited-absa Source: myjobmag
Summary drafted with AI assistance from the original advert. The advert itself is the authority — how we use AI.
At a glance
- AWS
- Azure
- Docker
- Kubernetes
- Agile / Scrum
Extracted automatically from the advert; confirm requirements on the original listing.
Job description
Before you apply
- Confirm the requirements and closing date on the original listing (myjobmag). SPANi lists vacancies from other sites and may not reflect last-minute changes.
- Legitimate employers do not charge application, registration or training fees.
- Don't send your ID or bank details before you have confirmed the employer is real.
Before you apply
- Read the full advert and confirm you meet the minimum requirements before applying.
- Tailor your CV headline and most recent experience to the job title and key skills.
- Note the closing date and reference number, and keep a copy of what you submit.
- Apply only through the employer or job board link — never pay to apply.