Skip to project content
Home Solutions Services Research Accountability Resources Book ∑ Scholar Company Community TEI

TERA · MATHEMATICAL REASONING · RESEARCH INITIATIVE

TERA Mathematical Reasoning Challenges

Precise problems. Complete reasoning. Evidence that can be examined.

TeraSystemsAI LLC is developing a private mathematical challenge and evaluation initiative to investigate where AI reasoning succeeds, fails, or requires further scrutiny.

We invite evaluation partners and mathematical professionals to help shape rigorous assessments across algebra, probability and statistics, topology, foundations of analysis, distribution theory, Fourier analysis, Sobolev spaces, and partial differential equations.

Current stage: development and pilot preparation.
A candidate problem collection has been prepared. Independent mathematical review and pilot validation are pending. We do not yet claim a certified benchmark, demonstrated model failure rates, or completed customer pilots.

What we aim to measure

A correct final answer does not by itself establish a valid proof. Our proposed evaluations examine both the result and the reasoning: assumptions, logical dependencies, treatment of edge cases, and the completeness of the argument.

A “stumper” is a challenge candidate, not a guarantee that a model will fail. Difficulty and failure claims must be supported by documented testing under specified conditions.

The proposed evaluation method

  1. Define the question. State one precise task with explicit assumptions, consistent notation, and an exact, verifiable answer.
  2. Review the mathematics. Prepare a complete reference solution and obtain independent review before admitting an item to scored testing.
  3. Fix the protocol. Agree on the model version, tool access, prompts, attempt budget, and scoring rubric before execution.
  4. Assess answers and proofs separately. Record errors, valid alternative solutions, uncertainty, and disputed judgments for adjudication.
  5. Report reproducible evidence. Document responses, conditions, limitations, and any retest exposure. Results apply to the tested configuration, not all mathematical capability.

The initial scope is exact-answer assessment and human review of natural-language mathematical arguments. Formal verification in Lean or another proof assistant requires a separately agreed scope.

Two ways to participate

Evaluation partners

For mathematical AI teams, model evaluation groups, research laboratories, and scientific reasoning organizations exploring a private evaluation pilot.

Begin with a discussion of the capability you need to measure, your current evaluation gaps, and the evidence your team would use. Scope, review capacity, confidentiality, timeline, and price are agreed before work begins.

Discuss a private pilot

Mathematical professionals

We welcome expressions of interest from mathematicians, statisticians, researchers, and experienced problem authors and reviewers.

Introduce your areas of expertise and whether you are interested in authorship or independent review. Any assignment, compensation, rights, and confidentiality terms must be agreed in writing. An inquiry does not guarantee an engagement.

Express professional interest

The TERA foundation

This initiative carries forward the founder-originated TERA philosophy that predates TeraSystemsAI LLC.

Trustworthiness
State what the evidence supports and make limitations explicit.
Efficiency
Use focused questions and proportionate evaluation effort.
Reliability
Review reference solutions and document repeatable testing conditions.
Accountability
Keep human review, decisions, and responsibility identifiable.

Confidentiality and intellectual property

The full collection and reference solutions are not published on this page. Access to demonstration or pilot materials is subject to agreed confidentiality and permitted-use terms.

Please do not include confidential problems, answer sets, employer or client materials, or third-party-owned content in an initial inquiry. Discuss your expertise and needs first; rights and submission terms must be resolved before material is accepted.

Challenge outcomes are not certification of an AI system’s safety, general competence, or regulatory compliance.

Start a conversation with TeraSystemsAI LLC

A Pennsylvania-based company developing research and AI assurance under TERA.

Project contact: business@terasystems.ai

Include your organization or professional background, the relevant mathematical domains, and whether your interest is a pilot or a contribution. No private problem submission is needed at this stage.

The inquiry links open your own email application. This page does not submit or send messages automatically.