Overview
The evaluation system automatically:- Analyzes claim extraction quality using industry-standard metrics
- Generates detailed evaluation reports in CSV format for local analysis
- Provides cloud dashboard access for advanced trace analysis
- Requires no custom code - evaluations run automatically in the pipeline
- Integrates with Confident AI for comprehensive evaluation tracking
DeepEval Framework
Confident AI Platform
Getting Started
API Key Setup
To access cloud dashboard features and detailed traces, you’ll need a Confident AI API key:Get your Confident AI API Key
Set your environment variable
Include in API requests
Basic Usage
Evaluations run automatically when you use CheckThat AI’s claim normalization features:Evaluation Metrics
The evaluation system assesses multiple dimensions of claim extraction quality:Accuracy Metrics
Accuracy Metrics
- Precision: How many extracted claims are actually valid claims
- Recall: How many valid claims were successfully identified
- F1-Score: Harmonic mean of precision and recall
- Semantic similarity between original text and extracted claims
- Factual consistency of extracted information
- Preservation of original meaning and context
Quality Metrics
Quality Metrics
- Coverage of all relevant claims in the source text
- Identification of implicit vs. explicit claims
- Handling of compound and nested claims
- Logical consistency of extracted claims
- Proper claim boundaries and segmentation
- Maintenance of causal relationships
Confidence Scoring
Confidence Scoring
- Model confidence in claim identification
- Uncertainty quantification for ambiguous cases
- Reliability scores for different claim types
- Assessement of how well-formed claims are for fact-checking
- Identification of claims requiring additional context
- Flagging of unverifiable or opinion-based statements
CSV Evaluation Reports
Receive detailed evaluation reports that you can save locally for analysis:Report Structure
Downloading Reports
Confident AI Dashboard Access
Access detailed traces and advanced analytics through the Confident AI cloud dashboard:Dashboard Features
Test Run Tracking
Evaluation Analytics
Model Comparison
Custom Metrics
Accessing the Dashboard
Ensure API Key is Set
CONFIDENT_API_KEY is included in your requests:Visit Confident AI Dashboard
View Your Test Runs
Analyze Performance
Dashboard Screenshots

Confident AI dashboard showing evaluation metrics and traces

Detailed test run view with claim extraction analysis
Integration Examples
Batch Evaluation Workflow
Best Practices
Evaluation Strategy
Evaluation Strategy
- Set up automated evaluation for production workflows
- Monitor evaluation trends over time
- Set quality thresholds and alerts for low-performance cases
- Choose evaluation metrics that align with your use case
- Balance accuracy, completeness, and processing speed
- Consider domain-specific evaluation criteria
Performance Optimization
Performance Optimization
- Process multiple claims together for better efficiency
- Use batch APIs when available for large-scale evaluation
- Implement proper rate limiting and error handling
- Monitor evaluation costs alongside regular API usage
- Use sampling for large datasets to control evaluation costs
- Balance evaluation frequency with budget constraints
Data Analysis
Data Analysis
- Regularly review CSV reports for quality trends
- Identify patterns in low-quality extractions
- Use insights to improve prompt engineering
- Leverage Confident AI dashboard for deep analysis
- Set up custom metrics for your specific domain
- Use trace analysis to debug extraction issues
FAQ
Do I need to write any code for evaluations?
Do I need to write any code for evaluations?
CONFIDENT_API_KEY to access dashboard features and detailed traces.How do I get CSV evaluation reports?
How do I get CSV evaluation reports?
What if I don't have a Confident AI API key?
What if I don't have a Confident AI API key?
CONFIDENT_API_KEY to access the cloud dashboard, detailed traces, and advanced analytics features.How often should I review evaluation reports?
How often should I review evaluation reports?

