Data Engineer & AI Engineer | DevOps — Crafting solid, working prototype/foundation, high-impact solutions.
- 🔍 Currently building: Data Quality Eval Engine, a data quality evaluation engine for LLM outputs and datasets (JSON/CSV/Parquet, CLI tooling, CI-tested)
- 🛠️ Focus areas: data engineering, dataset engineering, data pipeline reliability, LLM evaluation & AI observability
- 🌍 Open to: remote work roles, contract/freelance engagements, and open-source collaborations
- 📫 Reach me: Orig.Stuff.365@gmail.com · linkedin.com/in/Jab · [your portfolio site, if any]
Python · Pandas · PyArrow · Pytest · Git · GitHub Actions (CI/CD) · CLI Tooling · ETL Pipelines · JSON/CSV/Parquet · LLM Evaluation · Data Quality & Observability · MLOps
4 completed DataCamp career tracks (Data Analyst, Data Engineering, Associate Data Scientist in Python) - AI Fundamentals, AI Business Fundamentals + Azure Fundamentals (AZ-900) -- 84 courses, 1.3M+ XP, covering Databricks, Apache Airflow, ETL/ELT pipelines, and applied ML (NLP, XGBoost, reinforcement learning).
Data Quality Eval Engine
A CLI-driven engine that evaluates datasets for completeness and text quality across multiple formats. Fully tested (pytest), CI-verified on every push (GitHub Actions), and packaged for install via pip install -e ..
✅ Stable foundation, actively developed — core evaluation logic, CLI, multi-format loading (JSON/CSV/Parquet), 9 passing tests, CI verified on every push. 🚧 Not yet supported: database connections, large/streaming files, additional formats (Excel, XML). Planned as next milestones.
💬 Always happy to talk data quality, pipeline design, or LLM evaluation — feel free to open an issue or reach out directly.