B

o

n

j

o

u

r

;

 

I

'

m

S

u

g

u

m

a

r

a

n

B

a

l

a

s

u

b

r

a

m

a

n

i

y

a

n

A

I

/

M

L

E

n

g

i

n

e

e

r

Cloud, Data & AI Engineer with 7+ years of experience in production IT and systems. I build Python backend services, REST APIs, data pipelines, and cloud-native applications with AWS, Docker, Kubernetes, and CI/CD.

Since January 2026, I have been a Guest Professor / Intervenant at SKEMA Business School, teaching postgraduate Python, Machine Learning, NLP, Generative AI, Cloud Engineering, MLOps, and production systems.

</AboutMe>

I am a Cloud, Data & AI Engineer with 7+ years of experience across production IT and systems. My work spans Python backend development, REST APIs, data pipelines, cloud architecture, automation, and distributed systems.


I work with AWS, Docker, Kubernetes, CI/CD, PostgreSQL, IAM, data engineering, machine learning, and generative AI. Since January 2026, I have also taught postgraduate students at SKEMA Business School through practical courses in Python, Cloud, Data, AI, MLOps, and production-ready engineering practices.

Sugumaran Balasubramaniyan

</Experience>

Guest Professor / Intervenant — Python, Data, AI & Cloud

SKEMA Business School
Paris, France | Jan 2026 – Present

• Teach and mentor postgraduate students in Python, Machine Learning, NLP, Generative AI, Cloud Engineering, MLOps, CI/CD, data pipelines, and production systems.
• Design hands-on labs, exercises, demonstrations, and practical projects covering APIs, automation, testing, deployment, monitoring, and documentation.
• Review code and architecture, support debugging, and provide feedback on maintainability, performance, and production readiness.

AI / ML Engineer & Consultant — Cloud, Backend & Production Systems

Independent Consultant
Paris, France | Jan 2024 – Dec 2025

• Developed Python, FastAPI, REST API, PostgreSQL, Docker, and AWS services for business workflows and AI production systems.
• Designed architectures integrating APIs, databases, data pipelines, cloud components, and containerized services.
• Implemented Git, Docker, CI/CD, automated testing, versioning, monitoring, observability, logging, validation, and access-control practices.

Data Analyst / Data Engineer — Data Products & Automation

Joubert Associés
Paris, France | Jul 2023 – Dec 2023

• Developed automated pipelines and data models with Python and SQL across 50+ projects, reducing dashboard refresh time by 36%.
• Translated business needs into technical specifications, KPIs, data models, validation rules, and analytical solutions.

Process Associate — Cloud, Data Engineering & Machine Learning

Capgemini Technology Services
Chennai, India | Apr 2021 – Jul 2022

• Developed scalable pipelines and services with Python, AWS, Spark, SQL, REST APIs, AWS Glue, and S3, improving workflow speed by 45%.
• Industrialized workflows with Docker, Git, CI/CD, automated tests, AWS SageMaker, deployment, monitoring, and observability.
• Collaborated with Data, Engineering, Product, and business teams on Cloud/Data solutions for production environments, including IAM, access control, data quality, scalability, and reliability.

Process Executive — Automation, Data & Risk Systems

Infosys
Chennai, India | Mar 2020 – Apr 2021

• Developed Python and SQL automation solutions for risk and operations workflows supporting Citizens Bank.
• Implemented validation processes, quality controls, scenario testing, edge-case analysis, and failure identification.

Freelance Data Analyst — Python, Machine Learning & Analytics

Arimalytics
Puducherry, India | Jun 2018 – Feb 2020

• Developed Python forecasting and predictive-analytics solutions, improving forecast accuracy by 12% through feature engineering, benchmarking, evaluation, and optimization.

</Education>

Master of Science (MSc) — Data Science & Artificial Intelligence Strategy

emlyon business school
France | 2022 – 2024

• Exchange Program — McGill University, Montréal, Canada.

Master of Business Administration (MBA)

Pondicherry University
India | 2016 – 2018

Bachelor of Technology (B.Tech)

Pondicherry University
India | 2012 – 2016

</Certs>

AWS Machine Learning Engineer Associate

Amazon Web Services

• Proficient in designing, implementing, and deploying machine learning solutions on AWS.
• Expertise in SageMaker, feature engineering, and model optimization.
• Skilled in ML pipeline orchestration and automated model training workflows.

AWS Cloud Practitioner

Amazon Web Services

• Validated expertise in AWS cloud architecture and foundational services.
• Demonstrated knowledge of AWS pricing models and cost optimization strategies.
• Proficient in deploying scalable and secure cloud infrastructure on AWS.

Databricks AI Agents

Databricks Academy

• Mastered building autonomous AI agents using Databricks platform.
• Implemented LLM-based agents for complex task automation and reasoning.
• Expertise in prompt engineering and agentic workflow orchestration.

Snowflake Data Warehousing

Snowflake University

• Proficient in designing and managing cloud-based data warehouses using Snowflake.
• Expertise in data modeling, query optimization, and Snowflake governance.
• Skilled in data sharing and Snowflake collaboration features.

AWS GenAI Practitioner

Amazon Web Services

• Expert in building generative AI applications using AWS services.
• Proficient with Amazon Bedrock, SageMaker JumpStart, and generative AI tools.
• Skilled in prompt engineering and responsible AI practices.

Dataiku ML Practitioner

Dataiku Academy

• Proficient in end-to-end machine learning projects using Dataiku platform.
• Expertise in visual machine learning workflows and automated model selection.
• Skilled in model deployment and monitoring within Dataiku ecosystem.

Dataiku Developer

Dataiku Academy

• Expert in developing custom plugins and extensions for Dataiku platform.
• Skilled in Python development within Dataiku recipe and custom component frameworks.
• Proficient in integrating external APIs and data sources with Dataiku.

Atlassian Agile Project Management Professional

Atlassian Academy

• Expert in agile project management using Jira and Confluence platforms.
• Proficient in sprint planning, backlog management, and team collaboration workflows.
• Skilled in implementing agile methodologies and scaling agile practices across teams.

Microsoft Power BI Data Analyst Associate (PL-300)

Microsoft

• Certified in designing and building scalable data models, cleaning and transforming data, and enabling advanced analytic capabilities in Power BI.
• Proficient in DAX, Power Query, and publishing reports and dashboards for business decision-making.
• Applied directly in academic and professional contexts, including teaching BI at SKEMA Business School.

NVIDIA Certified Professional: Agentic AI

NVIDIA

• Certified in designing and deploying production-grade agentic AI systems using LLMs and tool-calling frameworks.
• Proficient in multi-agent orchestration, memory management, and RAG pipeline architectures.
• Skilled in building reliable, evaluable AI agents for enterprise environments.

NVIDIA Certified Associate: Generative AI LLMs

NVIDIA

• Certified in the fundamentals of large language models, transformer architectures, and generative AI techniques.
• Proficient in fine-tuning, prompt engineering, and deploying LLM-based solutions.
• Skilled in applying generative AI to real-world NLP and multimodal use cases.

Databricks AI Agent Fundamentals

Databricks Academy

• Proficient in building and evaluating AI agents using the Databricks platform and Unity Catalog.
• Skilled in integrating LLMs with tool use, retrieval augmentation, and structured outputs.
• Experienced in deploying agent workflows for enterprise-scale data and AI pipelines.

AWS Solutions Architect

Amazon Web Services

Databricks Platform Architect

Databricks

Databricks Advanced MLOps

Databricks

NVIDIA Accelerated Data Science Professional

NVIDIA

Dataiku ML/MLOps & Generative AI Practitioner

Dataiku

AWS Generative AI Practitioner

Amazon Web Services

</Languages>

English

Fluent

French

Professional

Tamil

Native

</Skills>

Tech Stack

  • Python
  • PyTorch
  • R
  • Scala
  • Azure-SQL
  • MySQL
  • Redis
  • PostgresSQL
  • GITHUB
  • HuggingFace
  • GIT
  • Anaconda
  • Apache-Spark
  • Apache-Airflow
  • Apache-Hadoop
  • Apache-Cassandra
  • Apache-Kafka
  • AWS
  • Azure
  • GCP
  • Power-BI
  • Tableau
  • NumPy
  • Pandas
  • Scikit-learn
  • Matplotlib
  • Plotly
  • Streamlit
  • Flask
  • Docker
  • Kubernetes
  • TensorFlow
  • HTML5
  • CSS3
  • JavaScript
  • React
  • Node.js
  • MongoDB
  • GraphQL
  • Confluence
  • Jira
  • Excel
  • FastAPI
  • OpenCV
  • Databricks
  • Snowflake
  • Dataiku
  • MLflow
  • LangChain
  • LangGraph
  • n8n

</Projects>

Enterprise Agentic RAG Platform

Enterprise Agentic RAG Platform

Production-grade RAG with PGVector HNSW hybrid retrieval (<20ms p95), autonomous multi-step agent orchestrator, and two-stage deterministic security guardrails (100% injection defense).

Technical Approach

Designed a dual-mode hybrid retrieval engine combining PostgreSQL 16 pgvector HNSW indexing (M=16, ef_construction=64) with BM25 lexical search via Reciprocal Rank Fusion (RRF). Built a bounded multi-step agent orchestrator dispatching vector search, cloud sizing calculations, and citation verification, wrapped in strict pre/post-execution security guardrails.

Key Results

  • Sub-20ms retrieval latency with 2,185+ QPS throughput on 1536-dim vectors
  • 100% factual grounding pass rate and 100% citation coverage across evaluation benchmarks
  • 100% defense rate against direct/indirect jailbreaks, SQLi, command injection, and obfuscated payloads
  • Full multi-format Document AI ingestion (MD, CSV, TSV, JSON, Scanned OCR)

Tech Stack

Python 3.12FastAPIPostgreSQL / PGVectorStreamlitSQLAlchemyDocker Compose
Patient Mortality Rate and Readmission Prediction

Patient Mortality & Readmission Prediction

ML fusion models (XGBoost + BERT) on large healthcare datasets achieving AUC-ROC of 0.81. Built scalable AWS pipelines (Glue, Athena, Lambda, SageMaker, Bedrock) reducing latency by 30%.

Technical Approach

Built a multimodal ML fusion system combining structured clinical data (XGBoost) with unstructured clinical notes (BERT). Deployed on AWS using a serverless architecture with event-driven inference pipelines.

Key Results

  • Fusion model achieved AUC-ROC of 0.81, outperforming single-modal baselines by 12%
  • Reduced inference latency by 30% using SageMaker endpoint optimization
  • Processed 100K+ patient records through AWS Glue ETL pipelines

Tech Stack

XGBoostBERTAWS SageMakerAWS LambdaAWS GlueAthena
Project 1

Customer Churn Prediction

Classification model (XGBoost, Random Forest) predicting customer churn on telecom data. Focused on high recall to enable targeted retention strategies before customers leave.

Technical Approach

Developed classification models (XGBoost, Random Forest, Logistic Regression) to predict customer churn on telecom subscription data. Engineered behavioral features from usage patterns and applied threshold tuning to maximize recall for at-risk customer identification.

Key Results

  • XGBoost achieved 85% recall, enabling proactive retention of 4 out of 5 churning customers
  • Contract type, tenure, and monthly charges were the top churn predictors
  • Delivered actionable retention segments for marketing team

Tech Stack

PythonXGBoostscikit-learnSeabornPandas
Project 2

Heart Stroke Prediction

Binary classification model predicting stroke risk from patient health indicators. Applied logistic regression, decision trees, and feature engineering on clinical data.

Technical Approach

Built binary classification models to predict stroke risk from patient health indicators. Applied logistic regression, decision trees, and ensemble methods with careful handling of class imbalance in clinical data.

Key Results

  • Random Forest achieved 94% recall on stroke cases
  • Age, hypertension, and glucose levels identified as top risk factors
  • Built SHAP-based explainability dashboard for clinical use

Tech Stack

Pythonscikit-learnSHAPPandasMatplotlib
Project 3

Sentiment Analyzer

NLP pipeline classifying sentiment from text reviews using BERT and traditional ML models. Deployed as a web application with real-time prediction capabilities.

Technical Approach

Built an NLP pipeline combining BERT-based transformer models with traditional ML (Naive Bayes, SVM) for sentiment classification. Fine-tuned DistilBERT on domain-specific text data and deployed as a Flask web application with real-time prediction API.

Key Results

  • BERT model achieved 92% accuracy vs 84% for baseline Naive Bayes
  • Sub-200ms inference time with model distillation
  • REST API handles 50+ concurrent requests

Tech Stack

BERTPyTorchFlaskHugging FaceNLTK
Project 4

Sleep Disorder Prediction

ML model predicting sleep disorders (insomnia, sleep apnea) from lifestyle and health metrics. Feature engineering on BMI, stress levels, physical activity, and sleep duration.

Technical Approach

Built multi-class classification models to predict sleep disorders (insomnia, sleep apnea, none) from lifestyle and health metrics. Engineered features from BMI, stress levels, physical activity, heart rate, and sleep duration data.

Key Results

  • Achieved 89% F1-score on multi-class sleep disorder classification
  • Stress level and daily step count emerged as the strongest predictors
  • Built feature importance visualizations for clinical interpretability

Tech Stack

Pythonscikit-learnXGBoostSeabornPandas
Project 5

Fraud Detection using R

IEEE-CIS fraud detection on 590K+ transactions using ensemble methods in R. Applied SMOTE to handle class imbalance and achieved strong AUC on the Kaggle benchmark.

Technical Approach

Applied ensemble methods (Random Forest, XGBoost, LightGBM) in R on 590K+ transactions from the IEEE-CIS Kaggle competition. Handled severe class imbalance (<1% fraud rate) using SMOTE and stratified sampling.

Key Results

  • Top 15% on IEEE-CIS Kaggle leaderboard
  • Achieved 0.91 AUC-ROC with stacked ensemble
  • Identified transaction amount and card verification as top fraud indicators

Tech Stack

RXGBoostSMOTEcarettidyverse
Project 6

Big Data Analysis using Databricks

Large-scale data analysis pipeline on Databricks using Apache Spark and Delta Lake. Processed millions of records to extract business insights using distributed computing.

Technical Approach

Designed and ran distributed data processing pipelines on Databricks using Apache Spark. Processed millions of records with Delta Lake for ACID-compliant data transformations and aggregations.

Key Results

  • Reduced query time by 60% using Delta Lake caching and Z-ordering
  • Built automated ETL pipeline processing 5M+ records in under 10 minutes
  • Created interactive dashboards directly in Databricks notebooks

Tech Stack

Apache SparkDelta LakeDatabricksSQLPython
Project 7

Medical Cost Prediction

Regression model predicting individual medical insurance costs from patient demographics. Explored feature interactions (BMI, age, smoking status) using Python and scikit-learn.

Technical Approach

Built regression models (Linear, Ridge, Random Forest, XGBoost) to predict individual medical insurance costs. Performed extensive feature engineering on BMI, age, smoking status, and region interactions.

Key Results

  • Identified smoking-BMI interaction as the strongest cost predictor
  • Achieved R² of 0.86 with XGBoost regression
  • Deployed as interactive Streamlit dashboard for what-if cost simulation

Tech Stack

Pythonscikit-learnXGBoostStreamlitPandas

</Reviews>

Sugumaran is a highly skilled engineer who has a deep understanding of the fundamental concepts and algorithms of Machine Learning, and their ability to implement these techniques in practical applications is remarkable. With their AWS Cloud Certified Practitioner certification, Sugumaran demonstrated a strong command of AWS services and tools, which they have utilized to design, build and deploy ML models.

MD

Mani Deva

Senior Software Engineer, Ivanti

I am happy to recommend Sugumaran for his exceptional skills in the field of Data Science and Business. Having worked closely with Sugumaran, I can confidently say that he is one of the top students I have had the pleasure of working with. He consistently demonstrated excellent technical skills in Data Science, and his ability to bridge the gap between technical and business aspects is highly valuable.

UP

Ulises Armando Ponce Sesma

AVP Special Credit Unit, Credit Risk | MSc Data Science & AI Strategy

I had the pleasure of working with Sugumaran on a project where he demonstrated his exceptional skills in data analysis and project management. He has a keen eye for detail and is skilled at identifying trends and patterns in complex datasets. His ability to communicate complex technical concepts in a clear and concise manner is a testament to his professionalism and dedication.

AK

ASHWATH KARTHIK

QA/QC Engineer | UPDA/MMUP Certified (Mechanical)

Sugumaran is one of the most technically rigorous ML engineers I have encountered. His ability to bridge research and production — from BERT fine-tuning to AWS SageMaker deployment — delivers real impact. A rare combination of deep technical skill and genuine passion for advancing the field.

Senior Data Science Leader, HealthTech

As an educator at a top business school, Sugumaran brought real-world AI/ML experience into the classroom. Students consistently rated his sessions among the most practical and engaging. His ability to explain complex ML concepts in clear, actionable terms is outstanding.

Academic Director, Business School

Worked alongside Sugumaran on a supply chain optimization project. His end-to-end ML pipeline design — from data engineering on AWS to model deployment — reduced forecasting errors by 21%. A rare blend of deep technical skill and sharp business acumen.

VP of Engineering, Supply Chain Technology

</Writing>

</Contact>