AI Data Engineer Β· 4+ Years

PhaniKumar S

Building pipelines at scale, with AI on top

Senior Data Engineer specializing in scalable data pipelines, real-time ingestion, and cloud-native architectures on Azure β€” extended with agentic RAG, Databricks Mosaic AI, and AI/BI Genie for natural-language, self-service data access. Currently at Amegy Bank, Houston.

4+
Years Experience
99.9%
SLA Adherence
60%
Processing Reduced
$80M
Revenue Impact
Houston, Texas, US
AVAILABLE FOR OPPORTUNITIES

Technical Stack

⚑ Stream Processing
Apache Kafka Azure Event Hubs Spark Streaming Delta Live Tables
☁️ Cloud Platforms
Azure Databricks Azure Data Factory Azure Data Lake Microsoft Fabric
πŸ”§ Data Engineering
Apache Spark Delta Lake DBT Apache Airflow Apache Avro
πŸ—„οΈ Storage & Warehouses
Azure Synapse ADLS Gen2 Parquet
πŸ’» Languages
Python SQL Spark SQL PySpark
πŸ€– AI & ML Engineering
Databricks Mosaic AI Agentic RAG LangChain Azure OpenAI Azure AI Search Chroma DB Vector Search Generative AI
πŸ“Š AI/BI & Governance
Databricks AI/BI Genie Power BI DAX Unity Catalog Azure RBAC

Experience

Amegy Bank
Sep 2025 – Present
Senior Data Engineer Β· Houston, TX
  • Orchestrated 25+ Azure Data Factory pipelines with dependency management, retry logic, and alerting via Azure Monitor, achieving 99.9% SLA adherence across all production pipelines.
  • Built a data validation and reconciliation framework using Azure Databricks with PySpark, improving transaction data accuracy to 99% and reducing financial reporting errors carrying an estimated $2M+ compliance risk exposure.
  • Developed automated failure detection via Azure Monitor and Logic Apps with real-time Microsoft Teams alerts, reducing incident response time.
  • Implemented Azure RBAC and Azure Active Directory (Entra ID) role-based access control policies, alongside Unity Catalog for centralized data governance, table/column-level access controls, and audit logging, to enforce least-privilege access on sensitive banking datasets in Azure Data Lake Storage Gen2.
  • Migrated full-load batch pipelines to incremental load architecture using watermarking and Delta Lake MERGE upserts on Azure Databricks, reducing data processing time and cutting annual cloud compute costs by $15K.
  • Automated CI/CD pipelines using Azure DevOps for data pipeline deployments, version control, and automated testing, improving code quality and reducing deployment time.
  • Built an agentic RAG POC using Databricks Mosaic AI Agent Framework, dynamically routing natural language queries between Genie (structured/SQL) and Mosaic AI Vector Search (semantic similarity) based on query intent, enabling accurate responses across banking data.
Wipro
oct 2021 – Dec 2023
Data Engineer Β· Hyderabad
  • Designed automated ADF pipelines with real-time triggers, enhancing operational reliability by 30%.
  • Optimized Spark jobs via partition pruning, caching, and Z-ordering on NestlΓ© sales data β€” improving query performance by 20%.
  • Developed automated pipeline failure detection using Azure Monitor + Teams alerts, achieving 50% faster incident response.
  • Implemented Delta Lake upserts with watermarking, reducing ETL processing time by 40%.
  • Supported NestlΓ©'s $80M YoY revenue growth in 2022 through optimized reporting pipelines.
Infosys
Mar 2020 – Sep 2021
Data Engineer Β· Hyderabad
  • Designed and deployed DBT models to ingest and transform insurance claims, policy, and provider data into the warehouse, processing 50M+ records daily and reducing reporting query time by 45%.
  • Developed complex transformation logic using Spark SQL, multi-level joins and window functions to validate claims adjudication, membership eligibility, and provider network data β€” improving data accuracy to 98% across 200+ end users.
  • Orchestrated insurance data pipelines using Apache Airflow DAGs β€” automating end-to-end claims ingestion, transformation, and loading into the warehouse for downstream analytics reporting.
  • Implemented DBT-based data lineage and access control policies to enforce data security compliance across sensitive policyholder datasets.
25+
Airflow DAGs Built
10+
Upstream Systems
40%
ETL Time Saved
3
Top-Tier Companies

Projects

Data Engineering
Flight Data Analysis Pipeline
  • End-to-end Azure Databricks pipeline using Databricks Autoloader for incremental ingestion (raw to bronze).
  • Delta Live Tables for silver-layer transformations and data quality enforcement.
  • Curated analytical views for top airlines and bookings-per-year trends.
Databricks Autoloader Delta Live Tables Medallion Architecture

Education

πŸŽ“
Avila University
Master of Science Β· Computer Science
Jan 2024 – Aug 2025
Kansas City, Missouri
Graduate Degree Computer Science MS
πŸ›οΈ
Kallam Haranadhareddy Institute of Technology
Bachelor of Technology Β· Computer Science
2016 – 2020
Andhra Pradesh, India
B.Tech Computer Science Engineering
πŸ… Certifications
Deloitte
AI Certified
Infosys
InfyTQ Certification

Get In Touch

Open to senior data engineering roles, consulting, and interesting pipeline challenges.

Available for Opportunities

Based in Houston, TX. Open to on-site, hybrid, or remote positions. Specializing in Azure-native data platforms, real-time streaming pipelines, and medallion architecture.

Full-time Contract Consulting Remote OK