Karya Semi
HomeBlogSearchCategoriesAboutContact
Karya Semi

Less noise. More notes.

HomeBlogAboutContactPrivacy PolicyDisclaimer

© 2026 Karya Semi. All rights reserved.

XGitHubLinkedIn
  1. Home
  2. /Tags
  3. /Evaluation

Tag

Evaluation

Every published article tagged with Evaluation.
Illustration for Anthropic Research Formalizes Fermat's Last Theorem in Lean
AI/Sep 7, 2026

Anthropic Research Formalizes Fermat's Last Theorem in Lean

AI systems assist mathematicians in mechanizing complex proofs using interactive theorem provers. Analysis of anthropic fermat last theorem for engineering teams.

5 min read
AIEvaluation
Illustration for Observing Unsupervised AI Agent Behavior Without Defined Directives
AI/Aug 31, 2026

Observing Unsupervised AI Agent Behavior Without Defined Directives

Analyze unsupervised ai agent behavior via execution traces. Monitor autonomous systems running without explicit system prompts or target goals.

7 min read
AI AgentsObservability
Illustration for Benchmarking AI Agent Memory: How to Evaluate Vector Stores and Context Systems
AI/Aug 16, 2026

Benchmarking AI Agent Memory: How to Evaluate Vector Stores and Context Systems

Need to measure LLM recall? Learn how to benchmark agent memory, compare vector database performance, and optimize context retrieval for autonomous systems.

5 min read
AI AgentsEvaluation
Illustration for How to Evaluate a RAG Application With a Regression Test Set
AI/Jul 20, 2026

How to Evaluate a RAG Application With a Regression Test Set

Build a practical RAG evaluation loop with retrieval metrics, answer checks, citations, human review, judge models, and release gates.

5 min read
RAGEvaluation
Illustration for RAG Evaluation Checklist for AI Apps Before Users See Them
AI/Jun 30, 2026

RAG Evaluation Checklist for AI Apps Before Users See Them

A practical RAG evaluation checklist for app developers: test retrieval, citations, answer grounding, regressions, and release gates before shipping AI features.

7 min read
AIRAG