FenoriTechnologies
FENORI RESEARCH

Understanding before intelligence.

Fenori Research examines how complex systems represent reality, make decisions, evaluate uncertainty and operate under consequence.

The work spans representation, state, algorithms, decision systems, applied AI, complex operations and the infrastructure that connects them.

SYSTEMS REPRESENTATION DECISIONS RELIABILITY AI OPERATIONS
R004 / RELIABILITY & EVALUATION · WORKING PAPER

From model performance to system evidence.

Model capability is one input into whether a system can be trusted in operation. It is not the same question.

Evaluation has to move upward from isolated model performance toward the behavior, context, failure modes and outcomes of the complete system. The benchmarks that changed the unit of analysis to real operational work report much lower numbers than the ones that test a model alone, and the standards bodies are moving the same direction.

The paper also takes the measurement problem seriously: the literature does not agree on how much observed inconsistency belongs to the model and how much belongs to the way we score it.

NIST · ICML 2025 · NAACL · EMNLP · PRINCETON · STANFORD HAI

Read the paper →
MODEL OUTPUT WAS THE ANSWER RIGHT MODEL IN A TASK DID IT COMPLETE THE TASK SYSTEM IN AN ENVIRONMENT DID IT HOLD UNDER REAL CONTEXT, ACCESS AND FAILURE OUTCOME OVER TIME DID THE ORGANIZATION GET WHAT IT NEEDED, REPEATEDLY
A schematic of the widening unit of analysis, not a measurement. Each step encloses the one before it, and each is answered by different evidence.

Six programs.

Fenori Research is organized around recurring technical questions rather than industries or product categories. Papers belong to more than one program where the question genuinely spans them.

01

Systems & Representation

Entities, relationships, identity, current and historical state, provenance, conflicting evidence, temporal change.

How should a system represent a reality that changes? Systems act against representations, not against the world. When a representation is incomplete, stale or structurally wrong, greater model capability does not reliably produce better decisions.

Time is the part most often demoted to metadata, and it is where representations tend to fail first.

Google Research and DeepMind built the Test of Time benchmark (ICLR 2025) from synthetic data so that temporal reasoning could be measured apart from memorized knowledge. Accuracy varied sharply with the structure of the temporal graph, the order facts were presented in, and the question asked.
02

Reliability & Evaluation

Repeated execution, perturbation, failure conditions, version change, calibration, task-level outcomes, system-level testing.

When is an intelligent system reliable enough to matter? Accuracy answers whether a system can succeed. Reliability asks whether an organization can depend on it, and the two have been coming apart.

Princeton's Towards a Science of AI Agent Reliability (ICML 2026) decomposes reliability into consistency, robustness, predictability and safety, and reports reliability improving at roughly half the rate of accuracy on a general agentic benchmark and about a seventh of the rate on a customer service benchmark. How much of that belongs to the models is itself contested: SCORE (NAACL 2025) measured accuracy moving up to ten points on MMLU-Pro under paraphrase alone, while Flaw or Artifact? (EMNLP 2025) argues much of that movement is an artifact of rigid answer matching rather than a property of the model. NIST's draft TEVV-Athlon framework points evaluation toward an organization's own objectives and real outcomes.
03

AI & Agents

Memory, task horizon, context compression, monitoring, long-running work, agent communication.

What changes when a system has to work over time rather than answer once? Long-running work introduces problems a single-turn evaluation never sees: what to remember, what to discard, when to wait, and how to recover a state that drifted while nobody was watching.

Agentic Memory (ACL 2026) treats storing, retrieving, updating and discarding as actions inside the agent's policy rather than as a retrieval database bolted alongside it. Microsoft Research's ACON (ICML 2026) reports cutting peak tokens by 26 to 54 percent through context compression while largely preserving task performance, which makes the discard decision an explicit engineering tradeoff. SentinelBench (Microsoft Research AI Frontiers) takes a different case entirely, 100 monitoring tasks across 10 environments where the world changes while the agent waits rather than because it acted.
04

Decision Systems

Evidence, policy, deterministic logic, probabilistic reasoning, human authority, escalation, reconstruction.

What should exist between information and consequential action? The interesting problem is rarely a model output on its own. It is the whole path by which information becomes a decision someone is accountable for, and the record that path leaves behind.

Part of that path is choosing the form of computation each step requires. Intelligence is not a synonym for probability, and some parts of a system become both safer and more useful when their behavior is deliberately constrained.

The EU AI Act requires that high-risk systems be built so competent people can oversee them, including the ability to interpret output, disregard it, override it or stop the system, with obligations that persist after deployment.
05

Physical & Operational Systems

Assets, people, events, dependencies, schedules, physical state, commercial obligations, changing conditions.

How should an operation be represented when no individual system explains the whole of it? Many operational failures are relationship failures. An event matters because of what it affects downstream, and the system holding the event is rarely the system holding the consequence.

NIST frames this as a measurement-science problem rather than a modeling one. Its work on credibility considerations for digital twins in manufacturing treats verification, validation and uncertainty quantification as what establishes whether a representation deserves to be trusted for its intended purpose, and NISTIR 8356 covers the security and trust considerations alongside it.
06

Governance & Evidence

Access boundaries, contextual integrity, information flow, identity, authorization, documentation, delegated authority.

What is a system allowed to know, what may it do with what it knows, and what has to survive so the decision can be examined later? An enterprise system does not need access to everything it could technically retrieve, and documentation written after deployment is not the same as a system built to preserve evidence while it runs.

Microsoft Research's CI-Work (ACL 2026) tested frontier systems on enterprise information flows and found privacy violation rates from 15.8 to 50.9 percent, with leakage reaching 26.7 percent, and a tension worth sitting with: higher task utility often came with more violations. The authors argue for context-centric architecture rather than model-centric scaling.

Research library.

Three tiers, marked plainly. Published work is available to read. Work in preparation is written but not yet released. The roadmap is the editorial plan, listed so the shape of the programme is legible rather than to suggest the papers exist.

Published

Current research

R004 Reliability & Evaluation From model performance to system evidence Primary sources  NIST · ICML 2025 · ICML 2026 · EMNLP · Stanford HAI Working paper
R005 Reliability & Evaluation Reliability is not accuracy Primary sources  Princeton, ICML 2026 In preparation
R006 Systems & Representation Representation before reasoning Primary sources  Google Research · ACL · NIST In preparation
R007 Governance & Evidence Context is part of the security boundary Primary sources  Microsoft Research, ACL 2026 In preparation
R008 Systems & Representation Time is part of state Primary sources  Google Research, ICLR 2025 In preparation

Roadmap

R009 Decision Systems · Governance Human oversight is an architecture problem Intended source base  EU AI Act · NIST AI RMF Not yet written
R010 Decision Systems When not to use AI Intended source base  NIST AI RMF · systems literature Not yet written
R011 Physical & Operational A digital representation has to earn credibility Intended source base  NIST digital twin VVUQ work Not yet written
R012 Physical & Operational Interoperability is not integration Intended source base  NIST · ISO 23247 Not yet written
R013 AI & Agents Memory is state management, not storage Intended source base  ACL 2026 · Microsoft Research Not yet written
R014 AI & Agents Longer tasks fail differently Intended source base  Microsoft Research, ICML 2026 Not yet written
R015 AI & Agents Compression changes what the system knows Intended source base  Microsoft Research, ICML 2026 Not yet written
R016 Reliability & Evaluation Continuous systems need continuous evaluation Intended source base  Microsoft Research · NIST TEVV Not yet written
R017 AI & Agents · Physical Monitoring is not just a long task Intended source base  Microsoft Research Not yet written
R018 Systems & Representation Natural language is not always the best internal representation Intended source base  ACL 2026 Findings Not yet written
R019 Reliability & Evaluation A model can be factual and still produce a bad decision Intended source base  NIST TEVV · Princeton Not yet written
R020 Governance & Evidence Documentation is part of the system Intended source base  NIST AI documentation work Not yet written
R021 Systems & Representation The current state is not the history Intended source base  Google Research · NIST Not yet written
R022 Physical & Operational Dependency is a form of information Intended source base  NIST systems-of-systems work Not yet written
R023 Decision Systems · Governance The boundary between automation and authority Intended source base  EU AI Act · NIST AI RMF Not yet written
EVIDENCE BASE

What we will cite, and what we will not.

Every factual claim, benchmark result, regulatory point or statement about the state of technology should trace to a source a CTO, regulator, engineer or researcher would recognize. Where the literature disagrees, we show the disagreement rather than the paper that flatters the argument.

PRIMARY STANDARDS AND LAW

NIST, EUR-Lex and the European Commission, CISA, OECD and official regulators.

PEER REVIEWED

ACL, NAACL, EMNLP, ICML, ICLR, NeurIPS, ACM and IEEE proceedings.

RESEARCH ORGANIZATIONS

Microsoft Research, Google Research and DeepMind, IBM Research, and university groups including Princeton, Stanford, MIT and Berkeley.

NOT CITED

Vendor marketing, consultancy trend reports, SEO articles and unsourced commentary. Industry sources only where they carry original research.

Each paper separates three kinds of statement: a fact drawn from a named source, a finding attributed to a named study and venue, and Fenori's own interpretation. The third is never presented as the first.

FENORI RESEARCH / MAP

The questions connect.

The areas are not separate subjects. Representation determines what decisions are possible, time determines what can be reconstructed, and context determines what a system may act on at all. Each node carries the work that sits under it.

REPRESENTATION R004 · R006 · P001 STATE R006 · P001 TIME R008 · B005 DECISIONS R003 · P003 · B004 ALGORITHMS R007 · P005 · B003 APPLIED AI R002 · R003 EVALUATION R005 · R003 CONTEXT R009 · P001 OPERATIONS R010 · P002 · P006
R references are Fenori research, P references are classes in the problem index, B references are published build architectures. The map is how the program is organized, not a claim about the field.

Research is useful when it changes what gets built.

Technology

How these principles become system architecture, and where control sits once they do.

Explore Technology →

Work

Where these patterns appear in real systems, and the build architectures they produced.

Explore Work →

Problems

The recurring classes of difficult problem the research is organized around.

Explore the problem index →

How we publish.

Fenori Research includes technical perspectives, engineering notes and working papers arising from our own work and from independent investigation.

We distinguish between published work and research currently in progress, and we mark which is which. Where a piece relies on external evidence, primary sources, standards and original research are cited directly so that a reader can check them.

Unless a publication states otherwise, Fenori Research is not peer reviewed. Figures are schematics of ideas rather than measurements, and are labeled as such.

Questions worth engineering.