Diverse Lynx is seeking an experienced AI Data Engineer to design, develop, and manage scalable data platforms that enable advanced analytics and AI solutions. The ideal candidate will build robust data pipelines and integrate AI/ML capabilities into enterprise data ecosystems.
Responsibilities:
- Design, develop, and maintain scalable ETL/ELT pipelines for structured and unstructured data
- Build and optimize data lakes, data warehouses, and AI-ready data platforms
- Develop ingestion, transformation, and orchestration frameworks using cloud-native technologies
- Prepare, cleanse, and engineer datasets for AI/ML and Generative AI workloads
- Integrate Large Language Models (LLMs), vector databases, embeddings, and RAG (Retrieval-Augmented Generation) pipelines into enterprise solutions
- Implement data governance, security, lineage, and quality controls
- Collaborate with Data Scientists, AI Engineers, Business Analysts, and Solution Architects
- Monitor, troubleshoot, and optimize data pipelines and platform performance
- Automate deployment, testing, and monitoring of data engineering workflows
- Create technical documentation and data dictionaries for enterprise data assets
Requirements:
- 15+ Years of experience as an AI Data Engineer
- Python
- AI/ML
- Retail domain experience is highly preferable
- ETL/ELT development
- Data Modeling (Star Schema, Snowflake Schema)
- Apache Spark
- Databricks
- Airflow
- Dataflow
- Informatica
- ADF
- Synapse
- or equivalent tools
- Relational & NoSQL Databases
- Data Warehousing concepts
- REST APIs and Microservices
- Git
- CI/CD
- DevOps practices
- Machine Learning fundamentals
- Data preparation for AI models
- Vector Databases (Pinecone, ChromaDB, FAISS)
- LLM Integration (OpenAI, Azure OpenAI, Gemini, Claude, etc.)
- RAG Architecture
- Embeddings and Semantic Search
- Prompt Engineering fundamentals