Summary

Trust is one of those words technology companies like to use when things get difficult. And sometimes when something is not really meant to become public.

Trust us with your data. Trust that our systems are secure. Trust that artificial intelligence is used responsibly. Trust that information is deleted when we say it is deleted. Trust that the model knows only what it is supposed to know.

We are building increasingly complex technical systems, connecting them to increasingly personal information, and simultaneously expecting people to place their trust precisely where they can see the least.

For us at FlameP, the principle is: evidence before trust.

We do not want people to have to trust FlameP because we seem likable, declare good intentions, or put a privacy seal on the website.

We want to build in a way that makes essential claims verifiable. Not everything, of course. Not every technical process. But what is crucial to the individual.

What information was used for my task?

What information was not used?

Which external provider was involved?

Which permission applied?

What was stored?

What was processed only temporarily?

When was a task complete?

And if we at FlameP claim to have made a social impact: Was the money merely budgeted, was it actually transferred, or can the impact itself also be substantiated?

1. The real problem is the asymmetry

Today, a digital system may know a great deal about its user.

The user knows remarkably little about the system.

They see an interface. They see an answer. Perhaps they also see the name of a model.

Numerous technical processes may be taking place underneath. Information is selected, supplemented, stored, sent to external systems, processed there, and then assembled again. Another model may check the first answer. A search engine may supply sources. A security system may classify content. Personal context that the user did not enter in this particular conversation may have been added.

The problem begins when people see only the final result and no longer have any reasonable way to understand how it came about.

Even for the operators of such systems, a complete overview is by no means a given. In its AI Risk Management Framework, the United States National Institute of Standards and Technology explicitly points out that AI systems can consist of many interdependent components and actors. Those responsible for one part of the system therefore often have neither complete visibility into nor complete control over other parts. Among other things, NIST recommends documented responsibilities, continuous risk management, monitoring of third-party systems in use, and traceable decisions across the entire life cycle. [1]

If even the organization building an AI application has to take active steps to understand its own technical supply chain, it is downright absurd simply to demand trust from the user.

2. Transparency is not yet evidence

Companies, including FlameP, like to respond to this with transparency.

We explain which providers we use.

We publish a privacy policy.

We describe our security measures.

That is necessary, but not sufficient.

Information about what is theoretically supposed to happen is different from evidence of what actually happened in a specific case.

A simple example:

A platform may state that it deletes personal data after a certain period.

That is a rule.

Evidence would be something else. It would have to demonstrate that the corresponding deletion rule was actually executed for a specific record.

The same applies to AI.

Key ideaWe use only the context necessary for your task.

That is a promise.

Key ideaFor this specific task, information areas A and B were used; area C remained excluded.

That is more verifiable.

Key ideaWe work with multiple AI providers and select the appropriate model.

That is a product description.

Key ideaThis task was processed on August 10 with Model X at Provider Y because it fell into this task category.

That is evidence of a specific process.

This is precisely the distinction that interests us.

Data protection law recognizes it as well. The GDPR’s accountability principle requires more than compliance with data protection principles. Controllers must also be able to demonstrate that compliance. The European Data Protection Board explicitly explains that, in practice, this means documenting data protection decisions and processes, among other things. [2] [3]

That is a different stance from “Trust us, we handle it properly.”

It says: If a claim is relevant, we must be able to provide evidence for it.

3. Regulation is already moving in this direction

The European AI Act also adopts this idea for certain AI systems.

For high-risk AI systems, the regulation requires, among other things, technical capabilities for the automatic logging of events. Such systems must also be designed with sufficient transparency to enable operators to interpret and use their outputs appropriately. Information about capabilities, performance limitations, intended use, and, where applicable, mechanisms for evaluating logs is likewise among the requirements. For high-risk systems, the AI Act also requires suitable means of human oversight. [4]

We should not draw the false conclusion that every AI product is subject to the same obligations. It is not. The AI Act deliberately uses different risk categories and different requirements.

We therefore find the underlying principle more interesting than the legal classification.

A technically complex system does not become more trustworthy because its manufacturer assures us that it is trustworthy.

Trustworthiness requires structures, documentation, logging, responsibilities, means of control, and known limitations.

And the ability to reconstruct errors later.

The NIST AI Risk Management Framework treats precisely these points as part of continuous governance. Documentation is intended to improve transparency, support human review, and strengthen accountability. The framework also calls for roles, risks, external components, human oversight, and system boundaries to be documented and regularly reviewed. [1]

The OECD has since gone one step further. Its Due Diligence Guidance for Responsible AI, published in 2026, connects traditional corporate due-diligence obligations with AI governance. Companies should identify risks and potential adverse impacts, implement measures, track their results, communicate about them, and enable remediation where appropriate. [6]

It is a shift away from assertion and toward verifiability.

4. But what is an ordinary person actually supposed to verify?

This is where things get complicated.

We could show a user a technical report containing 300 entries after every FlameP session.

Timestamps, API calls, model versions, token counts, hash values, storage operations, permission objects, network events.

That would be transparent, but not useful.

Transparency can itself become camouflage. Or, put differently: “Beat them with details.”

Anyone who floods people with enough information can subsequently claim to have disclosed everything. But nobody understood it.

We have long known this pattern from terms and conditions and cookie dialogs. Formally, much is explained. In practice, people click “Accept” because they do not want to invest half a working day before being allowed to use a website.

For FlameP, evidence before trust therefore does not mean maximum information.

It means relevant verifiability.

Users do not need to understand how a transformer works mathematically.

In ordinary use, they also do not need to inspect the entire internal routing algorithm.

But they should be able to get answers to a few very simple questions.

What did FlameP know for this task?

Where did information go?

What happened to it?

What did a machine decide?

What did a human decide?

What remains stored?

And what has now ended?

In our view, the true task of a good transparency layer is not to show everything. To us, transparency means showing the right things.

5. We therefore need different levels of evidence

We do not envision verifiability as a single enormous log file.

It requires at least three levels.

The first belongs to the user.

It must be comprehensible.

For a relevant task, a person should be able to determine, for example, which FlameP space was used, which context areas were authorized, which provider was involved, and whether processing was completed.

Not in developer language.

But roughly like this:

Key ideaContext used: Project FlameP and public author profile.
Key ideaNot used: personal conversations, health information, other projects.
Key ideaExternal processing: Provider X, Model Y.
Key ideaStorage by the provider: according to the terms applicable to this route.
Key ideaFlameP status: task completed, temporary authorizations closed.

This is not a complete audit.

It is an understandable receipt.

We need the second level ourselves.

It must allow what happened technically to be documented in substantially greater detail. Otherwise, we can neither investigate errors nor reconstruct security problems or check our own rules.

The third level becomes necessary where independent scrutiny is useful or required.

Auditors, supervisory authorities, or other authorized bodies may require deeper evidence that an ordinary user neither needs nor ought to see.

Three levels.

Comprehensibility for people.

Reconstruction for the operator.

Verifiability for independent scrutiny.

6. Yet this creates a new problem

Logging can itself become data collection.

This is one of those delightful contradictions that tend to disappear in a presentation.

We want to prove as effectively as possible what happened to data, so we log as much as possible. And suddenly we have built a new database that documents quite precisely what a person did and when.

That would be absurd.

Evidence before trust must therefore not mean retaining every process forever in personally identifiable form. Evidence, too, requires data minimization.

The GDPR identifies data minimization, purpose limitation, and storage limitation as core principles. Organizations should process personal information only to the extent and for as long as necessary for the respective purpose. [3]

For FlameP, this results in a demanding technical balancing act.

We must be able to demonstrate that something happened without unnecessarily retaining the entire content of what happened.

Sometimes an event can be logged without storing its content.

Sometimes it is enough to record that a particular type of context was authorized without copying the entire context into the audit log.

Sometimes a technical reference can prove that a particular version of a policy applied without storing all user data again.

And sometimes more detailed evidence will be required.

This cannot be solved with a single switch. It is hard architectural work.

7. AI answers make it even more difficult

Suppose FlameP clearly shows:

This model was used. These sources were retrieved. This context was authorized. These permissions applied.

All well and good, but of course that does not prove that the answer is correct.

This is an important limit of our principle. For many AI results, FlameP cannot prove that a statement is true. A language model can make an error despite a sound process.

Two models can be wrong. A source can be outdated. A scientific study can later be refuted. A business decision can fail despite excellent analysis.

Evidence before trust must therefore not itself become an illusion of absolute certainty.

In many cases, we can make the process more demonstrable. Which sources were used? What review was performed? Which uncertainties were identified? Was a second model used? Did a human approve the result? Which version of a piece of information was available at that time?

The result nevertheless remains a decision under uncertainty.

This is precisely why NIST distinguishes between the various characteristics of trustworthy AI. Transparency, explainability, privacy, security, reliability, and accountability are interconnected, but none alone guarantees that a system will act correctly in every situation. [1]

We therefore do not want to turn “Trust the AI” into a new “Trust the audit log.”

Evidence does not replace thinking, but it creates a better basis for it.

8. Providers must not disappear behind FlameP either

One of the most convenient options when building a multi-provider platform would be to make the technical complexity completely invisible.

The user talks to FlameP.

We select the model in the background.

Done. That would be convenient, but it would come at a cost.

FlameP would become a new central point of trust. The user would no longer be directly dependent on a single model provider, but would instead have to believe that we always do the right thing in the background.

That would be precisely the dependency we actually want to reduce.

Provider transparency is therefore part of evidence before trust for us.

This does not mean that someone must first make five technical decisions for every simple request. A good system may organize complexity.

But relevant decisions must remain reconstructable. Which provider was selected? Why was this route permitted? Which data class was allowed to reach this provider? Which rule version applied? Was there an alternative?

Was a particularly sensitive task handled differently from an inconsequential text draft?

In addition to transparency and explainability, the OECD AI Principles explicitly emphasize accountability and traceability as elements of responsible AI governance. [5]

To us, traceability is not a technical luxury.

Traceability determines whether provider independence truly creates independence or merely produces a new opaque intermediary.

9. Evidence before trust applies to us as well

Up to this point, one might get the impression that FlameP primarily wants to prove what other providers do with data.

If we say that personal context remains separate, our systems must actually enforce that separation.

If we claim that a FlameP space has access only to specific information, the entire personal history must not nevertheless be available in the background.

If a temporary permission has ended, it must be technically terminated.

If we claim that a task is complete, the temporary connections created for it should not simply continue indefinitely.

If we say that a particular process is not stored, our own logs must reflect that promise.

And if something does not work yet, we must not write as though it already does.

Product vision and product reality move at different speeds. A concept is approved in a workshop. The user interface already displays the new language. But part of the technical implementation is still missing in the backend.

The company begins talking about the product it wants to build as though it were already the product people use.

Capability before Claim.

What we claim must be supported by the system.

And if the system does not support it yet, we call it a plan, objective, or development. Not a feature.

10. The same applies to social impact

This point connects evidence before trust directly to our North Star of value flow.

Suppose FlameP decides to allocate 20.000 euros to a social program.

Then at least five different things could be meant. We have calculated that 20.000 euros would be earmarked for it under our rule. We have reserved this amount internally.

We made a binding commitment to a project. We actually transferred it.

Or we can additionally demonstrate that the money arrived and was used for the stated purpose.

These five states are not the same.

Corporate communications like to collapse them into one.

“20.000 euros for social impact.” Sounds wonderful. But afterward, no one knows what actually happened.

We want to distinguish these states, and certainly not because accounting is our favorite hobby.

We want to do so because social impact is particularly susceptible to well-intentioned self-deception.

The same applies to applications.

If we write that FlameP supports a particular language, it must be clear what “supports” means.

Can the system translate a few sentences? Can a native speaker actually work with it? Was the quality evaluated by people who speak that language? Does voice input work? Does it work only in a demo or in the real product?

11. Total transparency would still be a mistake

There is a romantic notion of transparency: if only everything were open, everything would be fine. But that is not true.

A security system should not publish every internal protection rule. Otherwise, it would hand potential attackers the instructions.

A user should not be able to see other users’ personal data merely because we promised transparency.

Nor do trade secrets and intellectual property disappear because a company wants to act responsibly.

And some technical information is not only incomprehensible but entirely irrelevant to the decision at hand. Evidence before trust therefore does not mean radical disclosure.

It means proportionate verifiability.

What must a particular actor know or be able to verify in order to assess a relevant claim?

The user needs a different view from a security assessor. A data protection authority needs a different view from a customer. A developer needs different information from an auditor.

For good reason, NIST works with roles, documented responsibilities, and different review processes rather than a single universal transparency requirement. [1]

12. Most importantly, however: trust will still be necessary

Our principle could now quite easily be used against us.

If everything is supposed to be proven, why have trust at all? Because no complex system functions without trust. No ordinary person will read the source code of every FlameP component. Nobody will personally inspect every provider’s data centers.

Even an independent auditor examines only certain areas at certain times.

Certifications can be flawed. Logs can be implemented incorrectly. Laws are broken. People make mistakes. There will always be a limit to what an individual user can verify personally. Our goal is therefore not to abolish trust.

Our goal is to reduce the area in which blind trust is necessary.

Today, much operates on the principle: Trust us first. Perhaps later you can verify whether we deserved it. We want to reverse that: We show as much relevant evidence as is reasonably possible. Trust can arise from that.

13. This also changes the distribution of power

In the end, evidence before trust is not a technical detail.

It is about power.

Whoever alone knows what a system has done possesses an informational advantage.

Whoever alone can decide which information becomes visible controls the narrative about it.

And whoever simultaneously owns the data, determines the rules, and writes the evidence of compliance with those rules is asking for a great deal of trust.

We will not be able to eliminate this asymmetry completely. FlameP operates the system. We therefore inevitably possess power. The relevant question is not whether power exists. For us, the relevant question is how we build mechanisms that constrain it.

Documentation constrains power. Traceable permissions constrain power. Separation of context constrains power. A visible provider constrains power. Independent assessments can constrain power.

And a corporate culture in which a claim requires evidence at least constrains the temptation to bring reality and marketing too close together.

The OECD now also explicitly connects responsible AI with a continuous process of risk identification, prevention, monitoring, communication, and remediation where appropriate. [6]

14. Conclusion: Our aspiration is simple. Implementing it will not be.

Today, we do not know how successfully we will achieve this in every area. Some evidence will be relatively straightforward technically.

Other evidence will be expensive. We will be able to present some information comprehensibly.

With other information, we will find that transparency and privacy collide.

Perhaps we will initially log things that we later have to handle differently.

Perhaps users will find that our transparency interface shows the wrong information.

Perhaps an external assessment will point out a blind spot that we did not see ourselves.

Evidence before trust does not mean that we will make no mistakes.

It means building our system, as far as possible, so that errors become visible, reconstructable, and correctable.

That is a more modest promise. And a more demanding one.

We do not want to say:

Key ideaFlameP is trustworthy.

That would be a claim about ourselves. We would rather say:

Key ideaHere is what FlameP did. Here are the relevant rules. Here are the limitations. And here is what can be verified.

Everyone can then decide for themselves how much trust arises from that. We believe this is the more honest order.

You do not have to trust FlameP. We have to earn that trust.

And wherever possible, our word should not be enough.

References

Key idea[1] National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework 1.0. NIST AI 100-1, January 26, 2023. The framework connects AI governance with documentation, responsibilities, monitoring, risk measurement, and traceability throughout a system’s life cycle. NIST is currently revising version 1.0.
Key idea[2] European Data Protection Board. Accountability. The GDPR’s accountability principle requires controllers not only to act in compliance with the rules, but also to be able to demonstrate compliance with data protection principles.
Key idea[3] European Union. Regulation (EU) 2016/679, particularly Article 5 on accountability, purpose limitation, data minimization, and storage limitation.
Key idea[4] European Union. Regulation (EU) 2024/1689, AI Act, particularly Articles 12 to 14. For high-risk AI systems, it includes requirements concerning logging, transparency, information for operators, and human oversight, among other things.
Key idea[5] OECD. OECD AI Principles, adopted in 2019 and updated in 2024. The principles address transparency, explainability, accountability, and traceability, among other things, as elements of trustworthy AI.
Key idea[6] OECD. OECD Due Diligence Guidance for Responsible AI. OECD Publishing, February 19, 2026. The Guidance applies corporate due-diligence obligations to the development and use of AI and calls for risk identification, prevention, outcome monitoring, communication, and remediation where appropriate, among other things.