Skip to main content
← Back to Blog
Security & DevOps6 min read

Encrypting AI Memory at Rest: A Practical Walkthrough

AES-256-GCM with a wrapped-DEK envelope and per-record HKDF key derivation, and the awkward part nobody mentions: encrypted rows break the vector index you were counting on.

Every encrypted memory row in Unimatrix carries exactly 106 bytes of cryptographic overhead: a format byte, a key-version byte, a wrapped tenant key (wrap IV, GCM tag, and wrapped ciphertext), a per-record nonce, and a data IV and GCM tag, concatenated in front of the ciphertext. That is the whole envelope format, and it is worth understanding byte by byte, because the interesting problems in encrypting AI memory are not in the cipher. They are in everything the cipher does not cover.

The layout

[Ver 1][KeyVer 1][WrapIV 12][WrapTag 16][WrappedDEK 32][RecordNonce 16][DataIV 12][DataTag 16][Ciphertext: N bytes]
   0    1         2          13            29             61              77          89      105    106            106+N

Each tenant holds a random 32-byte Data Encryption Key (DEK). The DEK is wrapped under the deployment root key (MASTER_ENCRYPTION_KEY) with AES-256-GCM, and the wrap IV, wrap tag, and wrapped DEK travel inline in the envelope. Then, per record, a fresh 16-byte nonce is generated and combined with the tenant DEK via HKDF-SHA256 to derive a record key used exactly once. The payload is sealed with AES-256-GCM under that record key, with the tenant id bound as additional authenticated data, and a fresh 12-byte data IV.

import { createCipheriv, createDecipheriv, hkdfSync, randomBytes } from 'node:crypto';

const IV_LEN = 12, TAG_LEN = 16, KEY_LEN = 32, NONCE_LEN = 16;

function wrapDek(dek, rootKey) {
  const iv = randomBytes(IV_LEN);
  const cipher = createCipheriv('aes-256-gcm', rootKey, iv);
  const wrapped = Buffer.concat([cipher.update(dek), cipher.final()]);
  return { iv, tag: cipher.getAuthTag(), wrapped };
}

export function seal(tenantDek, rootKey, tenantId, plaintext) {
  const { iv: wrapIv, tag: wrapTag, wrapped } = wrapDek(tenantDek, rootKey);
  const recordNonce = randomBytes(NONCE_LEN);
  const dataIv       = randomBytes(IV_LEN);
  const recordKey    = hkdfSync('sha256', tenantDek, recordNonce, `unimatrix/record/${tenantId}`, KEY_LEN);

  const cipher = createCipheriv('aes-256-gcm', recordKey, dataIv);
  cipher.setAAD(Buffer.from(tenantId, 'utf8'));
  const payload = Buffer.concat([cipher.update(plaintext, 'utf8'), cipher.final()]);
  const dataTag = cipher.getAuthTag();

  return Buffer.concat([
    Buffer.from([3, 1]), // format v3, key version 1
    wrapIv, wrapTag, wrapped, recordNonce, dataIv, dataTag, payload,
  ]);
}

The decipher.final() call is the one that matters. If the tag does not verify, it throws, and you get no plaintext at all. That is the difference between authenticated encryption and raw AES-CTR, where a flipped bit in storage silently becomes a flipped bit in the decrypted memory and your model happily reads corrupted context as fact. GCM makes tampering a hard error. The construction and its security bounds are specified in NIST SP 800-38D, which is also where the IV rules come from.

Why the per-record key is not paranoia

GCM with random 96-bit IVs has a birthday bound. Under a single key, the probability of an IV collision reaches roughly 2^-33 at around 2^32 encryptions, and an IV collision under the same key is catastrophic for GCM: it leaks the XOR of two plaintexts and, worse, allows forgery of the authentication key. Four billion memories is not an absurd number for a long-lived multi-tenant store.

Deriving a fresh key per record from a fresh nonce makes that bound irrelevant. Each key encrypts exactly one message, so the number of messages per key is one, and the collision analysis collapses. The cost is a single HKDF-SHA256 call per read and per write, which is a few SHA-256 blocks — effectively free next to the cipher operation itself. A heavyweight KDF (scrypt, PBKDF2) would add no security here: the tenant DEK already carries full 256-bit entropy, so stretching a lower-entropy password is unnecessary. Password derivation happens once, at the client, and only for passphrase-encrypted dashboard memories.

The part nobody puts in the launch post

Here is the uncomfortable structural fact about encrypting a memory store that also does semantic search. The memory text is ciphertext. The embedding vector next to it is not.

You cannot embed ciphertext. AES output is indistinguishable from random, so an embedding of it carries zero semantic signal, which means vector search over it returns nothing useful. The vector has to be computed from plaintext, before sealing, and then stored in a pgvector column in the clear so the ANN index can traverse it.

A 1024-dimensional float32 vector is not the plaintext, but it is a lossy projection of it, and lossy is not the same as safe. An attacker holding a stolen database dump without the master key can still do a great deal:

  • Query-by-guess. Embed candidate strings with the same public embedding model and rank them against the stored vectors. High cosine similarity to "quarterly revenue projections for the acquisition" tells you what a row is about without ever decrypting it.
  • Clustering. Group vectors and you recover the topic structure of a user's entire memory store, including how many distinct projects they have and which rows belong together.
  • Inversion. Published work on embedding inversion reconstructs substantial portions of short input text from its vector alone, given access to the same encoder. Short texts, which is exactly what memory entries are, are the easy case.

So the honest threat-model statement is narrower than "encrypted at rest." It is: encryption at rest protects memory content against disk theft, backup exfiltration, and snapshot leakage. It does not make the row opaque, because the vector beside it is a semantic side channel, and it does nothing at all against a compromise of the application tier, which by construction holds the master key in memory. Anyone claiming otherwise is selling you the cipher and not the system.

Partial mitigations exist and all of them cost something. You can encrypt the vector too and accept that semantic search now requires decrypting the whole column per query, which is fine at ten thousand rows and hopeless at ten million. You can quantize aggressively (binary or int8) to reduce the information content of each stored vector, which degrades both the attack and your recall. You can shard the ANN index per tenant so a partial dump only leaks one tenant's topic structure. None of these makes the leak zero.

Rotation is a schema decision, not a crypto decision

The scheme above wraps each tenant DEK under the root KEK. That makes rotating the root key a re-wrap operation, not a re-encryption of the content column: unwrap each tenant DEK with the old key and re-wrap it with the new one — two AES-GCM operations per tenant row and no KDF. The large ciphertext column is never touched. You still need a key-version byte in the stored envelope (that is what KeyVer is for) so that both keys are live during the rollout, and a migration that reads and rewrites the key rows of the tenants you are rotating.

The simpler alternative — deriving each record key directly from the master key with no wrapping — is clean and has one sharp edge: rotating the master means re-encrypting every row, because every record key is derived from it. Decrypt with the old master, derive a new salt, re-seal with the new master, under a key-version column so both masters stay live until the migration drains.

The rule of thumb: derive-from-master is the right call when your rotation cadence is measured in years or when the operator controls the whole deployment and can schedule downtime. Envelope with wrapped DEKs is the right call the moment rotation becomes routine, compliance-driven, or per-tenant. Both need a version byte in the stored blob from day one, because retrofitting versioning onto an unversioned format is a much worse migration than either rotation.

Unimatrix self-hosters supply their own MASTER_ENCRYPTION_KEY, which means the operator, not us, decides the rotation policy and owns the consequences. The full breakdown of what the ciphertext covers and what it does not is on the security page, including the vector side channel, because a threat model you cannot read is not a threat model.

Your AI remembers everything. Everywhere.

Unimatrix gives you a shared, durable memory layer across Claude Desktop, Cursor, ChatGPT, and Gemini. Setup in 2 minutes. Free and paid plans available.

Looking for developer resources?

Browse our catalog of 500+ tested AI prompt profiles covering DevOps, data modeling, agent behaviors, and API wrappers. Have a prompt to share? Submit your own for review by our librarian to be featured. Completely free, no registration required. Browse prompt libraries →

Keep reading