Skip to content
#

deepeval

Here are 149 public repositories matching this topic...

26-framework SDET monorepo: Playwright, Selenium, Cypress, Cucumber, Postman, FastAPI, Pact, AI/LLM eval (DeepEval), agentic AI (LangChain, LangGraph, DSPy, Claude), k6, flakiness detection, site monitoring, failure triage, QMS evidence (ISO 9001/SOC 2), dependency audit · GitHub Actions · DataDog · Terraform IaC

  • Updated Sep 4, 2026
  • Python

Advanced RAG pipelines for medical (HealthBench, MedCaseReasoning, MetaMedQA, PubMedQA) and financial (FinanceBench, Earnings Calls) QA. LangGraph orchestration + BAML structructed generation, Milvus Hybrid search (Dense + BM25 + RRF), three-layer Metadata Enrichment, Contextual AI instruction-following reranker, and DeepEval evaluation.

  • Updated Aug 20, 2026
  • Python

Advanced RAG pipeline optimization framework using DSPy. Implements modular RAG pipelines with Query-Rewriting, Sub-Query Decomposition, and Hybrid Search via Weaviate. Automates prompt tuning and few-shot selection using GEPA, SIMBA, MIPRO, COPRO, and BootstrapFewShot optimizers on datasets like FreshQA, HotpotQA, TriviaQA, Wikipedia and PubMedQA.

  • Updated Aug 20, 2026
  • Python

🚀 Production-ready modular RAG monorepo: Local LLM inference (vLLM) • Hybrid retrieval with Qdrant • Semantic caching • Docling document parsing • Cross-encoder reranking • DeepEval evaluation • Full observability with Langfuse • Open WebUI chat interface • OpenAI-compatible API • Fully Dockerized

  • Updated Jan 28, 2026
  • Python

Self-Reflective Question Answering for Biomedical Reasoning. GRPO fine-tuning via QLoRA & Unsloth with rewards for correctness, relevance, groundness, utility & XML structure. Structured think → answer → self-reflection with context grading, relevance assessment & groundness evaluation. DeepEval LLM-as-a-Judge (GEval, Faithfulness, Relevancy).

  • Updated Aug 6, 2026
  • Python

Production-grade prompt engineering curriculum — 7 modules covering foundations, reasoning, RAG, security, and evaluation. Model-agnostic (OpenAI/Anthropic/Ollama), runnable notebooks, mini-projects, Mermaid diagrams. Not prompt tips — prompt engineering as a discipline.

  • Updated Aug 20, 2026
  • Python

AI research assistant for prosthetics, orthotics & rehabilitation robotics papers. RAG pipeline with streaming citations, structured summaries, terminology extraction, and paper comparison. Built with Next.js 15, Claude, pgvector. Includes observability tracing and DeepEval evaluation pipeline.

  • Updated Aug 12, 2026
  • TypeScript

Production RAG system in Python: Haystack pipelines, FastAPI SSE streaming, Qdrant hybrid retrieval, OpenAI embeddings, DeepEval golden-set evaluation, and Langfuse tracing. Includes latency benchmarks (P50/P95 TTFT), retrieval failure-mode analysis, and chunking-strategy decision logs.

  • Updated May 26, 2026
  • Python

Add this topic to your repo

To associate your repository with the deepeval topic, visit your repo's landing page and select "manage topics."

Learn more