HomeSolutionsServicesResearchAccountabilityResourcesBook CompanyCommunityTEI

Open Source for Accountable AI

Research code, evaluation frameworks, and datasets published to enable scrutiny, reproducibility, and independent review. These resources support AI governance and safety evaluation, not autonomous deployment.

Featured Repositories

Reference implementations and evaluation frameworks for AI safety analysis and governance review

Safe Inference Framework

Reference inference framework demonstrating safety controls including guardrails, output filtering, and decision logging for audit discussions.

  • Guardrail and output filtering patterns
  • Decision logging with traceable audit trails
  • Configurable safety policy enforcement
  • Python, PyTorch, CUDA
Tradeoff Calculator

Interactive decision support tool for analyzing cost, quality, and latency tradeoffs during AI system planning and governance review.

  • Multi-axis tradeoff visualization
  • Architecture evaluation scoring
  • Documented decision export for governance records
  • TypeScript, React, D3.js
Bias Benchmark Suite

Benchmarking toolkit for evaluating bias and disparate impact across protected attributes and use cases in pre-deployment review.

  • Fairness evaluation across demographic groups
  • Regulatory documentation generation
  • Disparate impact quantification
  • Python, PyTorch, scikit-learn
ExplainML Interpretability

Unified interpretability toolkit providing SHAP, LIME, and attention visualization methods for explainability analysis and audit readiness.

  • Multiple explanation methods through a single interface
  • Audit evidence generation for compliance review
  • Model behavior documentation
  • Python, PyTorch, TensorFlow
PDF Secure

Document integrity evaluation toolkit combining cryptographic hashing and ML based tamper detection for forensic review and compliance workflows.

  • Cryptographic hash verification
  • ML based tamper detection
  • Traceable decision signals for legal contexts
  • Rust, Python bindings, WebAssembly
Uncertainty Quantification

Reference implementations for uncertainty estimation and calibration analysis, including conformal prediction, ensembles, and Bayesian methods.

  • Conformal prediction intervals
  • Model calibration assessment
  • Confidence alignment verification
  • Python, PyTorch, JAX

Open Datasets

Datasets published to support evaluation, benchmarking, and reproducible research

TERA-BIAS-21 CC BY 4.0

Bias evaluation dataset for NLP tasks across multiple protected categories. 250,000 annotated samples for research and evaluation.

  • Multi-category demographic coverage
  • Expert annotated with consensus labels
  • Reproducible evaluation protocols
DocTamper-Bench Apache 2.0

Document integrity and tamper detection benchmark containing 100,000 pristine and tampered documents for evaluation and benchmarking.

  • Paired pristine and tampered samples
  • Multiple tampering method categories
  • Forensic analysis ground truth labels
MedSafe-Eval Research Only

Healthcare AI safety evaluation dataset with 15,000 expert physician annotations for safety evaluation and governance review.

  • Clinical scenario safety annotations
  • Physician consensus labels
  • Risk severity classification
Uncertainty-Calib MIT

Multi-domain benchmark with 500,000 samples for uncertainty calibration and confidence assessment across evaluation and methodological research.

  • Cross-domain calibration evaluation
  • Confidence reliability measurement
  • Standardized comparison protocols
Open to contributors Build With Us on GitHub Browse issues labeled "good first issue" or "help wanted" to get started. Reviews prioritize correctness, documentation, and reproducibility. Explore Repositories →