AWISEE logo

AI Engineer (Data Guardrails & LLM Ingestion Pipelines)

AWISEE
1 day ago
Full-time
Remote
Worldwide
ML & AI Engineering

This is a remote position.

About the Role

We are looking for a highly skilled AI Engineer to design and build robust data ingestion, cleaning, validation, and LLM enhancement pipelines that power our AI applications. You will transform raw, unstructured data into high-quality, AI-ready datasets while implementing guardrails that ensure accuracy, consistency, and reliability.

Key Responsibilities

·         Design and develop scalable data ingestion pipelines for structured and unstructured data.

·         Build automated data cleaning, normalization, and preprocessing workflows.

·         Develop AI-powered enrichment pipelines using LLMs (OpenAI, Claude, Gemini, etc.).

·         Implement data quality validation and AI guardrails.

·         Develop prompt engineering workflows for data transformation.

·         Build document processing pipelines for PDFs, Word documents, CSVs, websites, and APIs.

·         Develop Retrieval-Augmented Generation (RAG) pipelines.

·         Create evaluation frameworks for LLM quality and accuracy.

·         Build ETL/ELT workflows for AI-ready datasets.

·         Integrate vector databases for semantic search.

·         Monitor pipeline performance, cost, latency, and data quality.

·         Collaborate with cross-functional teams to deliver production AI systems.

Required Technical Skills

Programming

·         Python (Expert)

·         SQL

·         Git

AI & LLMs

·         OpenAI API

·         Anthropic Claude API

·         Google Gemini API

·         Prompt Engineering

·         Function Calling

·         Structured Outputs

AI Frameworks

·         LangChain

·         LlamaIndex

·         DSPy (Preferred)

·         PydanticAI (Nice to Have)

Data Engineering

·         Pandas

·         Polars

·         ETL/ELT Pipelines

·         Apache Airflow (Preferred)

·         Data Validation Frameworks

Vector Databases

·         Pinecone

·         Weaviate

·         Qdrant

·         ChromaDB

·         FAISS

Cloud & Infrastructure

·         Docker

·         Kubernetes (Preferred)

·         AWS / Azure / GCP

·         Linux

Databases

·         PostgreSQL

·         MongoDB

·         Redis


Requirements

Preferred Qualifications

·         Experience building production-grade AI systems.

·         Strong understanding of RAG architectures.

·         Experience implementing AI guardrails and hallucination mitigation.

·         Experience with OCR and document parsing.

·         Experience with embedding models and semantic search.

·         Knowledge of data governance and security best practices.

Success Metrics

·         Build scalable ingestion pipelines.

·         Deliver automated data cleaning and LLM enhancement workflows.

·         Implement AI guardrails to improve output quality.

·         Develop evaluation pipelines for LLM performance.

·         Contribute to a production-ready AI platform.