R&D Evidence for ILR Application
Technical Innovation and Development Activities
1. Data Source Analysis and Matching R&D
50+ data feeds analyzed for optimal matching accuracy
Data Sources Analyzed
Matching Algorithm Performance Comparison
Key Finding: Hybrid approach combining rule-based matching with AI disambiguation achieved 96% accuracy, 29 percentage points higher than basic string matching. Thresholds tuned using 2,500+ labelled historic cases.
2. AI Model Experimentation
Testing multiple LLM providers and configurations
LLM Provider Performance Testing
| Model | Accuracy | Speed (ms) | Cost/1K | Traceability |
|---|---|---|---|---|
| GPT-3.5-turbo | 84% | 1,200 | $0.002 | ⚠️ Medium |
| GPT-4-turbo | 92% | 2,800 | $0.03 | ✓ High |
| Claude 3 Opus | 94% | 2,100 | $0.015 | ✓ High |
| ⭐ Claude 3.5 Sonnet (Constrained) | 97% | 1,600 | $0.008 | ✓ Very High |
| Gemini Pro | 88% | 1,900 | $0.0005 | ⚠️ Medium |
Hallucination Reduction
Innovation: Developed constrained prompting system requiring AI to cite specific source records for every claim. Reduced hallucinations by 91% (from 23% to 2%) while maintaining 97% accuracy. Tested on 1,500+ evaluation cases.
3. Workflow and Audit R&D
Testing different workflow architectures for optimal performance
Approach A: Human-First
Approach B: Rules-First
⭐ Approach C: AI-First (Hybrid)
Compliance & Audit Log Performance Optimization
Technical Breakthrough: AI-First hybrid workflow achieved optimal balance: 19x faster than human-first (2.5 vs 48 hours), 13% more accurate than rules-first (97% vs 84%), at only 21% of human-first cost ($95 vs $450).
4. Pilot Deployments and Iteration
Real-world testing with client organisations (Q1 2025 - Q2 2025)
Pilot Program Statistics
Error Rate Reduction Over Pilot Period
Key Improvements Implemented from Pilot Feedback
• PEP: +10% weight
• Old records: -20% impact
• Mobile-responsive
• Batch upload added
• PEP: senior review
• Clear cases: auto-approve
Validation Success: Pilot program validated technical approach and business model. Processing time 17x faster than industry average (2.8hrs vs 48hrs), error rate reduced by 87.5% through iterative refinement, achieving 4.8/5 satisfaction rating from 127 user reviews across 5 client organizations.