Karya Semi
HomeBlogSearchCategoriesAboutContact
Karya Semi

Less noise. More notes.

HomeBlogAboutContactPrivacy PolicyDisclaimer

© 2026 Karya Semi. All rights reserved.

XGitHubLinkedIn
  1. Home
  2. /Tags
  3. /Model Evaluation

Tag

Model Evaluation

Every published article tagged with Model Evaluation.
Illustration for Anthropic Research Formalizes Fermat's Last Theorem in Lean
AI/Sep 7, 2026

Anthropic Research Formalizes Fermat's Last Theorem in Lean

AI systems assist mathematicians in mechanizing complex proofs using interactive theorem provers. Analysis of anthropic fermat last theorem for engineering teams.

5 min read
AIEvaluation
Illustration for Anthropic Demonstrates Automated AI Alignment Improvement System
AI/Aug 31, 2026

Anthropic Demonstrates Automated AI Alignment Improvement System

New anthropic self improving ai system automates alignment. Fixes 10 misaligned behavior benchmarks. Baseline capabilities remain intact.

6 min read
AnthropicAlignment
Illustration for Mitigating Prompt Level Exploits and Cheating in Cyber AI Benchmarks
AI/Aug 22, 2026

Mitigating Prompt Level Exploits and Cheating in Cyber AI Benchmarks

Cyber AI benchmarks fail under exploit. Patch llm evaluation cheating prompt vulnerabilities to secure offensive security models against bypasses.

8 min read
Model EvaluationAI
Illustration for GLM-5.2 Token Costs Optimization: Writer Upgrades Post-Training Harness
AI/Aug 15, 2026

GLM-5.2 Token Costs Optimization: Writer Upgrades Post-Training Harness

Learn how a new validation harness optimizes post-training LLMs to reduce writer glm-5-2 token costs and maximize enterprise API efficiency.

5 min read
AIOpen Source
Illustration for GPT-5.6 Sol Preview: Why Model Upgrades Still Need Boring Evaluation
AI/Jun 29, 2026

GPT-5.6 Sol Preview: Why Model Upgrades Still Need Boring Evaluation

GPT-5.6 Sol may be stronger, but teams should test model upgrades with saved prompts, costs, latency, and failure cases before switching.

4 min read
GPT-5Model Evaluation