SERGAS Group
Head of Agentic Engineering and Transformation
to Present
Dubai, UAE · Full-time
Promoted from AI Engineer and Digital Transformation Manager (2024 to 2026). Owns AI engineering and transformation, delivering production systems with human-in-the-loop governance. Sole architect, now building the function and its team.
Head of Agentic Engineering and Transformation: sole architect of end-to-end digital transformation for a 35-year-old industrial enterprise. Built the company's data foundation from scratch: data lakes, medallion-architecture data marts (bronze/silver/gold), and temporal knowledge graphs (Graphiti) that unify decades of unstructured operational records into agent-accessible, ML-ready stores. All AI runs on-premise for data governance.
- Design and build multi-agent orchestrations: role-based agents, durable state, tool use, and verification loops.
- Prompt and context engineering with guardrails: chain-of-thought, ReAct, tool and function calling, and prompt-injection defence.
- Expose safe, narrow tool boundaries between models and live systems with MCP.
- Engineer agent memory and reliability: state persistence, multi-layer and reinforcement-learning memory, event-driven design, self-healing agents, and circuit breakers.
- Build retrieval-augmented generation systems: chunking, embeddings, hybrid retrieval, and reranking.
- Graph systems engineering: knowledge-graph modelling and graph-native retrieval over relationship-heavy domains.
- Embeddings and vector search: sentence-transformers, pgvector, LanceDB, Chroma, and cross-encoder reranking.
- NLP for legal and SOP reasoning as a domain skill of its own, modelling supersession, exceptions, and operationalisation so answers respect legal structure and currency.
- Run long-running training and inference loops on Apple MLX for on-device and on-premise workloads.
- Train and run inference on NVIDIA CUDA GPUs where the workload demands it.
- Parameter-efficient fine-tuning with LoRA and QLoRA on curated, PII-stripped domain data.
- On-premise model serving and quantization: Ollama, vLLM, llama.cpp, and GGUF.
- Edge and in-browser inference with WebGPU, ONNX Runtime, and WebLLM for private, low-latency inference close to the data.
- Computer vision and machine learning, extending into video and audio understanding (most recent work).
- Speech and multimodal pipelines: text-to-speech and speech-to-text.
- Build production applications across different tech stacks, chosen per environment and its constraints, rather than defaulting to one.