This is a remote position.
We are looking for a highly skilled AI Engineer to design and build robust data ingestion, cleaning, validation, and LLM enhancement pipelines that power our AI applications. You will transform raw, unstructured data into high-quality, AI-ready datasets while implementing guardrails that ensure accuracy, consistency, and reliability.
· Design and develop scalable data ingestion pipelines for structured and unstructured data.
· Build automated data cleaning, normalization, and preprocessing workflows.
· Develop AI-powered enrichment pipelines using LLMs (OpenAI, Claude, Gemini, etc.).
· Implement data quality validation and AI guardrails.
· Develop prompt engineering workflows for data transformation.
· Build document processing pipelines for PDFs, Word documents, CSVs, websites, and APIs.
· Develop Retrieval-Augmented Generation (RAG) pipelines.
· Create evaluation frameworks for LLM quality and accuracy.
· Build ETL/ELT workflows for AI-ready datasets.
· Integrate vector databases for semantic search.
· Monitor pipeline performance, cost, latency, and data quality.
· Collaborate with cross-functional teams to deliver production AI systems.
· Python (Expert)
· SQL
· Git
· OpenAI API
· Anthropic Claude API
· Google Gemini API
· Prompt Engineering
· Function Calling
· Structured Outputs
· LangChain
· LlamaIndex
· DSPy (Preferred)
· PydanticAI (Nice to Have)
· Pandas
· Polars
· ETL/ELT Pipelines
· Apache Airflow (Preferred)
· Data Validation Frameworks
· Pinecone
· Weaviate
· Qdrant
· ChromaDB
· FAISS
· Docker
· Kubernetes (Preferred)
· AWS / Azure / GCP
· Linux
· PostgreSQL
· MongoDB
· Redis
· Experience building production-grade AI systems.
· Strong understanding of RAG architectures.
· Experience implementing AI guardrails and hallucination mitigation.
· Experience with OCR and document parsing.
· Experience with embedding models and semantic search.
· Knowledge of data governance and security best practices.
· Build scalable ingestion pipelines.
· Deliver automated data cleaning and LLM enhancement workflows.
· Implement AI guardrails to improve output quality.
· Develop evaluation pipelines for LLM performance.
· Contribute to a production-ready AI platform.