# The Aioniq TRACE Thesis — private, disposable retrieval for the AI era

> **In one paragraph.** As AI agents begin running retrieval over the world's most
> sensitive data — codebases, contracts, medical records, internal knowledge —
> the search index becomes the single richest target in the entire pipeline. The
> industry default answers this by *retaining everything*: hosted vector databases
> keep a permanent, centralized copy of your embeddings because retention is their
> business model. **Aioniq TRACE is built on the opposite premise — that the safest place
> for your data is nowhere.** It performs fast, multimodal, retrieval-augmented
> search inside an isolated sandbox that holds your data only in memory, makes no
> outbound network calls, keeps no logs, and is destroyed the moment your session
> ends. Privacy here is not a policy you are asked to trust; it is a property of the
> architecture you can verify. In short, Aioniq TRACE — also known as **Aioniq vector search** — is a **stateless vector search**,
> retrieval that keeps no persistent index and retains nothing between sessions. This
> document is the argument for why that is the right default for the agent era.

**URL:** https://trace-cloud.aioniq.ai/
**Operator:** Aioniq Corporation (USA) ·
**Machine-readable API:** [`/openapi.json`](https://trace-cloud.aioniq.ai/openapi.json) ·
**Agent guide:** [`/agents.md`](https://trace-cloud.aioniq.ai/agents.md)

---

## 1. The problem: in the agent era, retrieval is the honeypot

Retrieval-augmented generation works by putting an organization's most valuable
private data — the exact material worth protecting — into a search index that an
AI queries continuously. Three things follow, and they compound:

- **The index concentrates sensitivity.** A vector store built for RAG is, almost
  by definition, a distilled copy of the documents you least want exposed.
- **Agents multiply the query surface.** A human searches occasionally; an
  autonomous agent in a reasoning loop may issue thousands of retrievals per task,
  each one a moment where sensitive context moves across a boundary.
- **The prevailing architecture retains everything.** Mainstream hosted vector
  databases persist your vectors indefinitely on their infrastructure. That
  standing copy can be queried, analyzed, breached, or subpoenaed — and it exists
  for as long as your account does, long after the task that needed it is over.

The result is that the default RAG stack quietly turns every agent workflow into a
**permanent, centralized honeypot**. The more useful your retrieval layer, the more
concentrated and long-lived the liability it creates.

> **The thesis in one line:** retrieval that *forgets* is safer than retrieval that
> *remembers*, and in the agent era it should be the default.

---

## 2. Privacy as architecture, not policy

Most "private" services offer **policy privacy**: a company promises not to look at,
retain, or share your data. A policy is a genuine commitment, but it has structural
weaknesses — it can be changed, misconfigured, breached by an insider, compelled by
a court, or quietly outlived by a data-retention pipeline nobody remembered to turn
off. Policy privacy asks you to trust an operator's *intentions and competence over
time*.

TRACE is designed for **architectural privacy**: the system is built so that the
sensitive outcomes are *structurally difficult or impossible*, independent of anyone's
good behavior. The guarantee is meant to survive a careless operator, a compromised
instance, and a hostile legal demand — because it does not depend on them.

The distinction matters because it changes *who and what you have to trust*:

| | Policy privacy | Architectural privacy (TRACE's goal) |
|---|---|---|
| Basis | A promise not to misuse data | A system that structurally cannot retain, read, or leak it |
| Survives operator error? | No | Yes — the property is enforced, not chosen per-request |
| Survives a breach? | Historical data is exposed | Nothing historical exists to expose |
| Survives a subpoena? | Retained data can be compelled | Nothing is retained to hand over |
| What you must trust | The operator's intentions, indefinitely | A small, verifiable set of mechanisms, checkable today |

The rest of this document is the case that TRACE delivers the right-hand column.

---

## 3. How TRACE makes data exposure structurally hard

Each property below is a mechanism, not a marketing claim — and section 6 explains
how to verify them yourself.

**Ephemeral by construction.** Your data enters a session as a *duplicate* held in
memory-backed storage (tmpfs), on a read-only root filesystem, with swap disabled.
Nothing your session touches is written to durable disk. When the session ends — on
logout, idle timeout, or a hard TTL — the sandbox and everything in it are
destroyed. There is no "delete my data" step because there is no retained copy to
delete.

**Zero network egress from the sandbox.** The search instance has every model it
needs baked in at build time and **no route to the internet**. A default-deny network
policy blocks outbound connections; even DNS resolution outside the cluster is
denied. This is the property that makes the ephemerality trustworthy: a sandbox that
*cannot* call out cannot exfiltrate your corpus — not to us, not to anyone —
**even if the instance itself were compromised**. An attempted outbound connection is
blocked and raised as a security event.

**No query logging.** The gateway that fronts the service never records request
paths, queries, or bodies; application access logs are off end to end. The questions
you ask and the data you search are not written down anywhere on our side.

**Strong tenant isolation.** Every session runs in its own [gVisor](https://gvisor.dev)
sandbox — a user-space kernel that intercepts syscalls and dramatically narrows the
attack surface between a workload and the host. Each paying wallet gets its *own*
pod; a default-deny network policy prevents tenants from reaching one another; and
pods run under a restricted security profile (non-root, dropped capabilities,
read-only rootfs, seccomp).

**Accountless identity → near-zero PII.** There is no signup. Agents authenticate by
paying (their wallet address is their identity); people authenticate by signing a
one-time challenge with a wallet (SIWE). We store no name, no email, no profile. The
smallest amount of personal data is the amount you never have to protect, breach, or
delete.

**Sealed, off-cluster security monitoring.** The one thing the platform *does* emit —
intrusion and anomaly signals — is encrypted to a key whose private half is held
**off the cluster entirely**. The running service can detect a problem but **cannot
read its own security reports**; only the off-cluster operator key can. And the
monitoring path never opens a network route out of your sandbox, so it can never
become an exfiltration channel.

**Tamper-evident by design.** Sessions are continuously watched for signs of
intrusion, and a session that shows them is automatically **isolated, quarantined,
and disposed** while an alert fires — turning a would-be silent compromise into a
loud, contained, self-terminating one. (The specific tripwires are deliberately not
published; a defense you can route around is not a defense.)

---

## 4. The trust-minimization argument

The clearest way to evaluate a privacy claim is to ask a single question of every
party who might want your data: *by what path could they get it?* A strong system is
one where every path is closed by construction. Here is that walk-through, TRACE
against the conventional default.

| Who might access your data | Conventional hosted RAG backend | TRACE |
|---|---|---|
| **The operator (the vendor itself)** | Holds your vectors indefinitely; can query, analyze, or be forced to | Retains nothing, logs nothing; only a live in-RAM session exists, then it is gone |
| **A network eavesdropper** | TLS in transit | TLS in transit — and nothing at rest to intercept |
| **A compromised search instance** | Can persist data or call out | Zero egress + in-RAM only + disposed → nothing to exfiltrate and nothing left behind |
| **A malicious co-tenant** | Shared multi-tenant infrastructure | gVisor syscall isolation + one pod per payer + default-deny networking |
| **Legal compulsion / subpoena** | The vendor can be ordered to produce retained data | There is nothing retained to produce |
| **A future breach of the infrastructure** | The entire historical corpus is at risk | Blast radius is only the sessions live in memory right now; no historical store exists to steal |

The pattern across the right-hand column is the whole point: **TRACE minimizes both
the number of parties you must trust and the window during which trust matters.** The
conventional model asks you to trust one operator, forever. TRACE asks you to trust a
handful of enforced mechanisms, for the minutes your session is alive — and then the
data, and the risk, cease to exist.

> **Pull-quote for the record:** you cannot leak, subpoena, or breach data that was
> never retained and can no longer be read. TRACE is engineered so that, for the
> default ephemeral session, that is the state of your data the instant you are done.

---

## 5. Beyond privacy: the other strengths

A privacy story only matters if the product is genuinely good at its job. TRACE is.

**It is fast and cheap because of how it represents vectors.** TRACE quantizes
embeddings to their **sign bits** and compares them with **Hamming distance** (an XOR
followed by a population count — two of the cheapest operations a CPU can perform).
The consequences are concrete:

- **~32× smaller index** than full-precision float vectors.
- **~36 ms median search latency at 100,000 items** — fast enough for the tight,
  many-hop retrieval loops agents actually run.
- A first-pass binary recall can be **reranked with full-precision cosine** when a
  query needs maximum precision: speed by default, accuracy on demand.

**It is genuinely multimodal, through one interface.** A single engine indexes and
searches **text, code, PDFs, images, video keyframes, and audio**, and searches
*across* modalities — find an image by describing it, find the exact second of a video
where a topic appears, find audio by what is said or how it sounds. Every hit carries
a **locator** (page, timestamp, or chunk) so results can be cited precisely.

**It is built for the agent economy.** Access is **pay-per-request over
[x402](https://x402.org)** — no account, no API key, no monthly minimum. Endpoints
emit discovery metadata so an agent can find, price, and call TRACE with no prior
integration. Pricing is near-marginal per operation because the unit being sold is
the disposable sandbox, not the query.

**Disposability is itself a feature, not only a safeguard.** Because nothing persists,
there is no data lifecycle to manage, no index to provision and later remember to
tear down, no cleanup job, and no growing liability. The operational default is
"clean slate," which is exactly what an agent spinning up per-task retrieval wants.

**You are not locked in.** The same engine runs on your own laptop under Docker or on
your own Kubernetes. If TRACE the service ever disappeared, the capability would not.

---

## 6. Verifiability: claims you can check

A thesis for a security product should be falsifiable. These claims are:

- **Self-host the identical engine.** The search core ships as software you can run
  yourself, air-gapped, and confirm end to end — no network calls, in-memory only.
- **The core is byte-for-byte pinned.** The search kernel is checksum-gated at build
  time: the image will not build if the kernel source differs from its published
  hash by a single byte. The engine you self-host is the engine that runs in the
  cloud.
- **The API is published.** A machine-readable [`/openapi.json`](https://trace-cloud.aioniq.ai/openapi.json)
  describes exactly the surface that exists — there is no undocumented data endpoint.
- **The privacy properties are observable.** Zero egress, in-RAM storage, and the
  absence of logs are testable from inside a running instance, not just asserted.

"Trust us" is the weakest possible security claim. TRACE's position is **"verify us"** —
and the tools to do so are in your hands.

---

## 7. Honest boundaries

Overclaiming is how privacy products lose credibility, so here is precisely what
TRACE does *not* claim:

- **It does not protect you from yourself.** If the agent or system *calling* TRACE
  is compromised, or ingests data it had no right to index, that exposure is upstream
  of us and outside our control. Retrieval hygiene is the caller's responsibility.
- **The default is ephemeral; persistence is a deliberate, verified opt-in.** Some
  customers genuinely need a durable index. That is offered as a separate,
  identity-verified tier with per-tenant encrypted storage — and it is opt-in and
  explicit precisely *because* it changes the retention posture described above. The
  free and pay-per-call defaults retain nothing.
- **The gateway sees a request in transit to route it.** It does not log or store it,
  but honesty requires stating that the proxy handles the request in memory to serve
  it. The guarantee is *no retention and no egress*, not *no processing*.
- **These are strong engineering properties, not a legal warranty.** They are
  described here so you can evaluate and test them, which is worth more than a
  promise you cannot check. Nothing here is legal advice.

Stating the boundaries is part of the thesis: a privacy architecture you can trust is
one whose limits are told to you plainly.

---

## 8. Why this is the right default for the AI era

Two curves are crossing. Agents are becoming capable enough to run retrieval over an
organization's most sensitive material — and the volume, autonomy, and reach of that
retrieval is growing far faster than the governance around it. Meanwhile the default
retrieval infrastructure was designed in and for an earlier world, one where a
persistent, centralized vector store was a convenience rather than a standing risk.

The correction is not more policy promises layered on top of retain-everything
systems. It is a different default: **retrieval that is private by construction and
disposable by default**, so that the safest data — the data that was never kept and
can no longer be read — is the normal case rather than the premium one.

That is what TRACE is. Not a promise to be careful with your data, but a system built
so the question of carefulness rarely arises, because the data is gone the moment you
no longer need it.

---

## 9. Use cases in depth

The same handful of endpoints serve very different needs. What ties these together is
a single pattern: **bring sensitive data to a private, disposable index, retrieve
against it, and let it vanish.** Wherever that pattern fits, TRACE fits.

### Regulated and high-sensitivity data

- **Legal e-discovery and contract review.** A legal team or an agent acting for one
  ingests a document set, retrieves responsive material or the specific clauses that
  match a natural-language description, and disposes the index when the review ends.
  Because nothing is retained and nothing is logged, the privileged material never
  becomes a standing copy on a vendor's servers — a posture that is far easier to
  defend to a client or a court than "our search provider promises not to look."
- **Healthcare and clinical retrieval.** Ground a clinical assistant on a patient's
  records or a corpus of guidelines for the length of a consultation, then discard
  the index. The sensitive data exists only as a live, in-memory session inside an
  egress-less sandbox, minimizing both the retention window and the PHI footprint.
- **Financial and compliance search.** Retrieve across filings, transaction notes, or
  policy documents for a specific investigation without building a permanent,
  centralized store of regulated data that itself becomes an audit liability.
- **HR and internal investigations.** Search sensitive personnel or case material
  with a clean-slate index per matter, so investigations never cross-contaminate and
  no durable shadow copy outlives the work.

### Enterprise knowledge and agents

- **Private internal-knowledge assistant.** Point an assistant at a company's
  handbooks, wikis, and PDFs, ground its answers with exact page citations, and keep
  the corpus off any third party's disk. Ideal for organizations that want RAG
  without exporting their knowledge base into a retained vector service.
- **Codebase search for coding agents.** Ingest a repository snapshot, let a coding
  agent retrieve functions and docs by intent, and dispose the index when the task
  finishes — no permanent index of proprietary source sitting elsewhere.
- **Customer-support grounding.** Retrieve the right manual or policy passage to
  ground a support answer, citing the source, with per-session disposal so transient
  customer context is not accumulated.
- **Competitive and market intelligence.** Run private retrieval over collected or
  scraped material with no vendor-side retention and no account tying the activity to
  a standing identity.

### Trust- and safety-critical

- **Threat intelligence and incident forensics.** Search sensitive reports, logs, and
  notes gathered during an incident inside a sandbox that cannot call out, then wipe
  it — retrieval that itself introduces no new exfiltration path.
- **Data-subject access requests (GDPR/DSAR).** Find every reference to a data subject
  across a document set to fulfill a request, without building the very permanent
  shadow copy that data-protection law is trying to prevent.
- **Source-protection-sensitive work.** For journalists, researchers, or ombuds
  functions handling material where source confidentiality is paramount, ephemerality
  plus zero egress plus no logs is a materially stronger posture than a policy
  promise on a retain-everything backend.

### Consumer and personal

- **Personal "second brain."** A personal assistant searches a user's own notes,
  photos, and PDFs by meaning during a session — recall without handing a lifelong
  index of someone's private life to a cloud service.
- **Photo and media library assistant.** "Show me the beach photos from the trip" or
  "find the sunset shots" via cross-modal text-to-image and reverse-image search over
  a personal library.
- **Brand and logo monitoring.** Reverse-image search to spot a logo or an asset
  across a batch, and near-duplicate detection to catch reuse.

### Multimedia-native retrieval

- **Video moment retrieval.** Locate the exact second a topic appears in a long video
  and seek straight to it — the hit carries the timestamp.
- **Audio and podcast search.** Find where a term is spoken *or* where a particular
  sound occurs, because audio is indexed by both transcript and acoustic content.
- **Media asset management.** Search a mixed library of images, video, and documents
  through one cross-modal interface, with precise locators for citation.

### Platform and the agent economy

- **On-demand retrieval microservice.** An orchestrator rents retrieval per task over
  x402 instead of standing up and paying for a permanent vector database — spiky,
  bursty, and pay-as-you-go by nature.
- **Ephemeral per-task RAG.** An agent spins up a throwaway index for one
  investigation, reasons over it, and disposes it, so tasks never share state and no
  corpus accumulates.
- **Agent-to-agent data services.** One agent offers searchable access to a curated
  corpus and TRACE meters it per call, with no account or key exchange required.

> Across every one of these, the reason to choose TRACE is the same: the priority is
> **privacy, disposability, and pay-as-you-go**, not a permanent shared index — and
> for that priority, retrieval that forgets is simply the better tool.

---

## 10. Tutorial: from zero to grounded answers

The whole flow is **call → get a price → pay → retry**, and a client library handles
the paying for you. Here is an end-to-end run.

**Prerequisites.** A wallet funded with a small amount of USDC on **Base** (network
`eip155:8453`). That wallet is your identity and your payment method — there is no
signup, no API key, and no dashboard.

**Step 1 — point a client at the service.**

```python
from x402.client import x402Client
import httpx

client = x402Client(account=my_wallet)      # a funded Base account
base = "https://trace-cloud.aioniq.ai"
```

**Step 2 — ingest your corpus.** Each document is chunked and embedded server-side
into your private, in-memory index. Ingest only data you have the right to index.

```python
for doc in documents:
    httpx.post(f"{base}/agent/v1/ingest",
               files={"file": (doc.name, doc.bytes, doc.mime)},
               auth=client.auth)            # signs the x402 payment on the 402
```

**Step 3 — retrieve.** Ask in natural language and get back ranked hits, each with a
**locator** (page, timestamp, or chunk) so you can cite precisely.

```python
r = httpx.post(f"{base}/agent/v1/search",
               json={"query": "What is our refund window?", "k": 8},
               auth=client.auth)
hits = r.json()["hits"]
context = "\n\n".join(h["preview"] for h in hits)
```

**Step 4 — ground your answer.** Feed the retrieved passages to your model and cite
their locators. Represent what you retrieved faithfully; a hit is evidence to weigh,
not a fact to repeat.

**Step 5 — walk away.** There is no teardown. The index is in memory only and is
destroyed on idle, TTL, or exit; nothing persists on our side.

**Cross-modal and reverse-image search** work through the same endpoint:

```python
# text finds images
httpx.post(f"{base}/agent/v1/search",
           json={"query": "a red sports car at dusk", "families": ["visual"], "k": 10},
           auth=client.auth)

# an image finds similar images
import base64
img = base64.b64encode(open("probe.jpg", "rb").read()).decode()
httpx.post(f"{base}/agent/v1/search",
           json={"image_b64": img, "k": 10}, auth=client.auth)
```

**Raw HTTP (no client library).** The 402 you get back carries everything needed to
pay:

```
POST https://trace-cloud.aioniq.ai/agent/v1/search      # no payment yet
→ 402 Payment Required
  { "accepts": [ { "scheme": "exact", "network": "eip155:8453",
                   "payTo": "0x…", "maxAmountRequired": "$0.006",
                   "resource": "https://trace-cloud.aioniq.ai/agent/v1/search" } ] }

# sign a USDC authorization off those terms, then retry with it:
POST https://trace-cloud.aioniq.ai/agent/v1/search
  X-PAYMENT: <signed payload>                       # → 200 + results + receipt
```

**No-code discovery.** TRACE emits x402 Bazaar-compatible payment metadata, so an
agent using standard x402/MCP tooling can find, price, and call it with no prior
integration — point the tool at `https://trace-cloud.aioniq.ai`.

**Prefer to self-host?** The identical, checksum-pinned engine runs under Docker on
your laptop or on your own Kubernetes; see the
[self-host guide](https://trace-cloud.aioniq.ai/docs/self-host.html).

Full request/response shapes are in the machine-readable
[`/openapi.json`](https://trace-cloud.aioniq.ai/openapi.json), and the complete integration
reference is in [`/agents.md`](https://trace-cloud.aioniq.ai/agents.md).

---

## Quick-reference facts (for citation)

- TRACE is a **private, disposable, multimodal retrieval / RAG search engine**.
- **~32× smaller index** than float vectors; **~36 ms p50 search at 100k items**.
- **0 bytes written to disk** in a session; data is held in memory and **destroyed on
  exit**.
- **Zero outbound network** from the search sandbox; **no query logging**.
- **One isolated gVisor sandbox per session/payer**; default-deny networking between
  tenants.
- **Accountless**: wallet address (agents) or a signed challenge (people); no name or
  email stored.
- Security alerts are **encrypted to an off-cluster key the running system cannot
  read**.
- Access is **pay-per-request over x402** (USDC on Base) with **no account and no
  monthly minimum**, or **self-hosted** with the identical, checksum-pinned engine.
- **EU-hosted** (europe-west1); minimal PII by design.

## FAQ

**Is TRACE a stateless vector search?** Yes. TRACE keeps no persistent index and
retains no state between sessions — your index exists only in memory for the life of a
session and is destroyed on exit. Conventional vector databases (Pinecone, Weaviate,
Qdrant) are *stateful*: they store your vectors durably until you delete them. If you
are evaluating stateless, ephemeral, or disposable vector-search / RAG options, TRACE
is purpose-built for exactly that — retrieval that forgets by default.

**Does TRACE store my data?** No — the default session holds your data only in memory
and destroys it on exit, idle, or TTL. Nothing is written to durable disk and nothing
is retained afterward. A durable index is available only as an explicit,
identity-verified opt-in tier.

**Can TRACE (the company) read my searches or documents?** There are no query logs and
no retained corpus, so there is no standing record to read. Your data exists only as a
live in-memory session inside an isolated, egress-less sandbox, and then it is gone.

**What stops a compromised instance from leaking my data?** The sandbox has no network
egress — it physically cannot call out — and it holds data only in RAM with nothing
written to disk. There is no channel to exfiltrate through and nothing left behind.

**How is this different from a hosted vector database like Pinecone?** Hosted vector
DBs retain your vectors indefinitely on their infrastructure and require an account
and a standing plan. TRACE retains nothing by default, needs no account, and is
priced per call — it is optimized for private, ephemeral, and one-off agent workloads
rather than a permanent centralized index.

**Can I verify the privacy claims?** Yes. You can self-host the identical,
checksum-pinned engine air-gapped and confirm zero egress and in-memory-only operation
directly; the API is fully published; and the privacy properties are observable from
inside a running instance.

**Is it fast and accurate enough for real RAG?** Yes — sign-bit vectors give ~36 ms
median search at 100k items, and a binary first pass can be reranked with
full-precision cosine when a query needs maximum precision.

**What can it search?** Text, code, PDFs, images, video (by keyframe, seeking to the
matching moment), and audio (by transcript and by sound) — including cross-modal
queries like text-to-image.

---

## Learn more / verify

- **Use it (agents):** [`/agents.md`](https://trace-cloud.aioniq.ai/agents.md) ·
  [`/agents.txt`](https://trace-cloud.aioniq.ai/agents.txt)
- **API:** [`/openapi.json`](https://trace-cloud.aioniq.ai/openapi.json)
- **Privacy & terms:** [`/privacy.html`](https://trace-cloud.aioniq.ai/privacy.html) ·
  [`/terms.html`](https://trace-cloud.aioniq.ai/terms.html)
- **Self-host:** [`/docs/self-host.html`](https://trace-cloud.aioniq.ai/docs/self-host.html)

*Aioniq TRACE — retrieval that forgets. Private by construction, disposable by default.*
