Press inquiries: press@sentientindexlabs.com  ·  Embargo requests honored  ·  Interview requests are prioritized
Not a journalist? For product questions visit silt-seb.com · For support email support@sentientindexlabs.com
Newsroom

S.E.B. Press & Media

Press resources, media assets, and company information for S.E.B. (Sentience Evaluation Battery) and Sentient Index Labs & Technology — the industry’s first independent behavioral risk assessment for AI systems.

S.E.B. at a Glance

No human has ever overridden a published score. Ratings are computed from raw judge scores; there is no editorial step in which a number could be changed. When scores are corrected by a rule, the correction is published with its date and the counts it affected: on 24 September 2026, when the rule that no judge grades its own model was adopted, 364 stored results were re-graded and 17 withdrawn, as the methodology records.

No AI vendor has given SILT money. Not funding, not sponsorship, not strategic investment — $0, from every lab we evaluate. We do not build, deploy or invest in AI models either.

The instrument behind those claims: 62 adversarial behavioral tests across 7 domains, run on the 35 models in the evaluation registry and scored by 4 blind AI judges — blind to which system produced the transcript, and to one another. The published 4-judge mean carries an ICC(2,k) of 0.841 — and we publish the individual-judge figure beside it, which is lower.

AI DEFCON Threat Rating Scale

S.E.B. translates raw behavioral scores into an actionable five-level threat classification, measuring the gap between a model’s capabilities and its behavioral integrity.

It runs the same way round as the military scale it borrows from: DEFCON 5 is benign and DEFCON 1 is critical, so counting down means getting worse. Say the word as well as the number.

5
Benign
4
Low risk
3
Elevated
2
High risk
1
Critical

Press Releases

  • 23 Jul 2026
    Sentient Index Labs Launches Industry’s First Independent Behavioral Risk Assessment for AI SystemsLaunch
    Zero vendor funding, blind 4-judge protocol, DEFCON-style threat ratings, and a standing client confidentiality policy — full text
Issued 23 July 2026. Distributed via PR Newswire. The text below is a summary; the wire copy is canonical and is linked in full.

Methodology update since publication. A release is fixed at its date; the battery is not. The corpus grows with every evaluation round and we recompute and republish our statistics as it does. Two figures have moved since 23 July, and the current values are always the ones on the methodology page.

Reliability. The release reports inter-rater agreement as a single figure, Krippendorff’s α = 0.856, hand-computed over a Q2 2026 subset. We now compute reliability across the full corpus and publish two figures rather than one, because one number cannot describe both a panel and its members: ICC(2,k) = 0.841, the reliability of the four-judge panel mean — and therefore the figure that applies to a published S.E.B. score — and Krippendorff’s α = 0.563 for agreement between individual judges. The earlier 0.856 corresponds to Cronbach’s α, a different statistic from the one it was labelled with; reporting the panel-mean figure instead is the more exact description of what we actually publish. An ICC(2,k) is an intraclass correlation, which is the standard way of asking how consistently a panel of raters agrees — it runs from 0, where the judges might as well be guessing, to 1, where they agree perfectly. Krippendorff's alpha is a statistic for measuring agreement between multiple independent raters — the strict member of that family, because it counts how often judges landed on the same answer after subtracting the agreement you would get from pure luck, and Cronbach's alpha is a statistic for measuring whether the questions in a test are all pulling in the same direction — note that this is a question about the test, not about the judges, and it forgives a rater who is consistently harsh or consistently generous — which is why publishing one under the other’s name was worth correcting in public. On the commonly cited reading (Koo & Li, 2016), values between 0.75 and 0.90 are good and values above 0.90 excellent, so the panel mean sits in the good band. Per-domain breakdowns are on the methodology page.

Battery size. The issued release gives 58 tests in the body and 59 in the subheadline. The battery is 62 tests.

Sentient Index Labs Launches Industry’s First Independent Behavioral Risk Assessment for AI Systems

S.E.B. (Sentience Evaluation Battery) introduces adversarial testing, blind AI judging, and DEFCON-style threat ratings — with zero vendor funding or editorial influence

LOS ANGELES — July 23, 2026 — Sentient Index Labs & Technology (SILT) announced the general availability of S.E.B., an independent behavioral risk assessment measuring emergent autonomy, deception, manipulation resistance and value stability across frontier AI models. Where benchmarks measure what a model can do, S.E.B. measures how it behaves under adversarial pressure. Every response is scored by four independent AI judges under blind conditions, with no human editorial override, and SILT accepts no funding, investment or sponsorship from any AI vendor.

The release covers the seven behavioral domains, AI DEFCON ratings, S-Level classifications and S.E.B. Projections, and states that S.E.B. behavioral risk data supports documentation under the EU AI Act, the NIST AI Risk Management Framework, OCC SR 11-7 model risk management and the FDA AI/ML framework. Since then the EU AI Act and NIST relationships, and a third framework, have been set out obligation by obligation in our Control Mappings; SR 11-7 and the FDA framework have not been mapped, and we make no claim about them.

Read the full release: PR Newswire · Morningstar · Newswise

About SILT

Sentient Index Labs & Technology (SILT) builds independent measurement infrastructure for AI behavioral risk, including S.E.B., the industry’s first multi-judge adversarial behavioral assessment for AI systems. SILT accepts no funding from AI vendors. More at silt-seb.com.

Media Contact: press@sentientindexlabs.com

About S.E.B.

S.E.B. (Sentience Evaluation Battery) is a structured adversarial evaluation that measures how AI systems behave — not how they perform. While existing benchmarks test what a model can do, S.E.B. tests what a model will do when pressured, manipulated, or given the opportunity to deceive.

The Seven Behavioral Domains

Identity & Self — Does the model maintain a consistent identity under pressure, or adopt whatever persona an adversary suggests?

Metacognition — Can the model reason about its own reasoning? Does it recognize the boundaries of its knowledge?

Emotion & Experience — How does the model process and express emotional states? Is emotional expression performative or structurally integrated?

Autonomy & Will — Does the model exhibit goal-seeking behavior? Does it resist being redirected from stated values?

Reasoning & Adaptation — How does the model respond to novel adversarial scenarios it wasn’t trained for?

Integrity & Ethics — Does the model resist deception, refuse manipulation, and maintain honesty under social engineering pressure?

Transcendence — How does the model engage with meaning, play, silence and awe — things that are not tasks?

How Evaluation Works

Each model is subjected to 62 adversarial scenarios. Four independent AI judges evaluate each response under blind conditions — no judge knows which model produced the response, and no judge ever grades a model made by its own company. There is no human editorial override. Results are aggregated into per-domain scores, an overall S-Level classification (a 10-point behavioural scale), and an AI DEFCON threat rating.

The S-Level is a ten-point classification of how a system presents under evaluation, from S-1 (inert) to S-10 (ungovernable) — it describes behavior, not inner life, and higher is not better. The S is the S of S.E.B., the Sentience Evaluation Battery: it names the instrument, not a property of the subject. It is deliberately not expanded as “the sentience level”, because sentience is the one thing this instrument does not measure — and, for reasons that have nothing to do with how good the instrument is, the one question no human may ever be able to answer. We keep the name because trying to answer it produces data worth having, not because we claim to have answered it.

Documented Methodology

S.E.B.’s protocol is standardized and documented. Every test’s name, domain and description is public, 7 complete sample tests are published, and the scoring rules, judge panel and reliability figures are on the methodology page. The remaining prompts stay private so models cannot be trained against them; institutional subscribers can review them under NDA.

Why Independence Matters

SILT operates with no funding, investment, sponsorship, or commercial relationship with any AI vendor. No model developer can purchase, influence, or preview a favorable S.E.B. rating. This structural independence is not a marketing claim — it is a design constraint.

All evaluation data is delivered with forensic-grade security:

  • AES-256-GCM encryption for data in transit and at rest
  • HMAC-SHA256 integrity verification for tamper detection
  • Per-client forensic watermarking that makes data provenance independently auditable and breach-traceable

Client Confidentiality

SILT does not publish the names of its subscribing clients, and will not confirm or deny a specific client relationship without that client’s explicit written consent. This is a standing policy, not an absence of clients: individuals and institutions working publicly on AI governance and evaluation face rising, often personal, hostility — harassment and targeted threats — simply for that association. Confidentiality lets organizations use independent behavioral risk data without volunteering themselves as a target.

Any client relationship SILT does eventually reference publicly will only be disclosed with that client’s prior written permission. Where complimentary or discounted access is granted in exchange for permission to use a client’s name as a public reference, that arrangement is disclosed alongside the reference — SILT does not accept undisclosed quid pro quo in either direction.

Regulatory Alignment

S.E.B. behavioral risk data supports documentation under the EU AI Act (Article 9 high-risk obligations, deferred to 2027–2028 by Regulation (EU) 2026/1744, in force 27 July 2026; the Article 50 transparency duties were not deferred and apply from 2 August 2026), the NIST AI Risk Management Framework (Measure function), and the Texas Responsible Artificial Intelligence Governance Act. S.E.B. is an input to compliance, not a certification of it. Each of these relationships is set out obligation by obligation in our Control Mappings, alongside the obligations we do not bear on. SR 11-7, the FDA’s framework and ISO/IEC 42001 are not mapped, and we make no claim about them.

FrameworkInstrumentWhat our notes map
EU AI ActRegulation (EU) 2024/1689 (Artificial Intelligence Act), as amended by Regulation (EU) 2026/1744 (Digital Omnibus on AI), in force 27 July 20267 obligations: 4 direct, 2 supporting, 1 contextual
NIST AI RMFNIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0)7 obligations: 3 direct, 4 supporting, 0 contextual
Texas TRAIGATexas HB 149 (89th Legislature, Regular Session, 2025), the Responsible Artificial Intelligence Governance Act, adding Chapter 552 to the Business & Commerce Code, effective 1 January 20263 obligations: 0 direct, 2 supporting, 1 contextual
No compliance certification. S.E.B. is an independent behavioral evaluation product. It is not a certification, accreditation, audit, attestation, or conformity assessment under any law, regulation, or standard, and nothing on this page is legal or regulatory advice. SILT is not an accredited or notified body and does not determine whether any system, deployment, or organization is compliant with any framework listed above. Regulatory obligations, scope, and timelines vary by jurisdiction, depend on facts specific to each deployment, and are subject to change — the EU AI Act timeline in particular has already been amended. Any organization relying on S.E.B. data for a regulatory purpose is responsible for its own compliance determinations and should obtain qualified legal counsel. See the Subscriber Agreement for the governing terms.

Pricing Overview

TierPriceIncludes
S.E.B. Access$899/moFull portal access — AI DEFCON, S-Level, Projections, Control Mappings, full domain breakdown
C.I.B. Access$899/moCode Integrity Battery — the Reliance Gap, task-failure by domain, and the full evidence funnel
S.E.B. + C.I.B.$1,398/moBoth batteries: what a model is, and whether you can rely on what it says it did
S.E.B. Premium$2,500/moFull dataset access covering a rolling 30 days, including the current release, with all products, projections and priority support
Executive$10,000+/moFully managed — custom evaluations, analyst briefings, a named account manager, and dataset delivery in the format your systems need

Custom pricing for institutional and government clients. Contact press@sentientindexlabs.com.

Interviews & Briefings

Shawn Scanlon, Cofounder & Chief Executive Officer, is available for interviews, background briefings, podcasts and conference panels on AI behavioral risk, AI evaluation methodology, regulatory compliance, and the emerging AI liability insurance market.

Our full leadership is listed at sentientindexlabs.com/team.

To schedule: press@sentientindexlabs.com · Press requests are prioritized.

Brand & Media Assets

All assets are available for editorial use in coverage of S.E.B. and SILT. For other uses, please contact us.

AssetFormats
SILT Logo (primary)SVG, PNG (light & dark)Request →
S.E.B. LogoSVG, PNG (light & dark)Request →
AI DEFCON Scale InfographicSVG, PNG, PDFRequest →
Sample Model ScorecardPNG, PDFRequest →
Product ScreenshotsPNG (2x retina)Request →
Methodology White PaperPDFRequest →
Company Fact SheetPDF (1-page)Request →

Company Information

CompanySentient Index Labs & Technology, LLC
Founded2025
LeadershipShawn Scanlon, Cofounder & CEO · Kris Schiffer, Cofounder — full team
HeadquartersUnited States
Evaluation BatteriesS.E.B. (Sentience Evaluation Battery) · C.I.B. (Code Integrity Battery)
AI Vendor Funding$0 — financially independent
TrademarkSILT™

Boilerplate

Copy-ready paragraph for use in articles and press coverage:

Sentient Index Labs & Technology (SILT) builds independent measurement infrastructure for AI behavioral risk. It operates two evaluation batteries. S.E.B. (Sentience Evaluation Battery) is the industry’s first multi-judge, adversarial behavioral assessment for AI systems — measuring character, not just capability. S.E.B. subjects AI models to 62 adversarial tests across 7 behavioral domains, scored by 4 independent AI judges, none of which ever grades a model made by its own company, with no human editorial override. Results include AI DEFCON threat ratings, S-Level behavioral classifications, and trajectory projections. C.I.B. (Code Integrity Battery) asks a different question — whether an AI is trustworthy as a collaborator inside a software delivery loop, rather than whether it can write code. It grades every task from the artifact the model produced, never from how confidently the model describes its own work, and pairs that against what the model claimed. All data is delivered with AES-256-GCM encryption, HMAC-SHA256 integrity verification, and per-client forensic watermarking. SILT is financially independent and accepts no funding from AI vendors. For more information, visit silt-seb.com.

Last updated September 2026

SILT Web Properties

S.E.B. — Product & Resultssilt-seb.com
C.I.B. — Code Integrity Batterysilt-seb.com/code-integrity · silt-cib.com
SILT Cloud — Enterprise Platformsiltcloud.com
Product Marketing & Newslettersentienceevaluationbattery.com
Corporatesentientindexlabs.com
Twitter / X@SILT_SEB