Presenting the demonstration vault
The counts and sample questions below describe the corpus that ships with SkyKeep, which is what this deployment installs — SKYKEEP_DEMO_CORPUS_URL is unset, so nothing has replaced it. If that changes, this note changes with it.
An ordered list of things to SHOW, each with its click path and the point it makes. The vault's thesis is compartmented reach, so the spine is beats 2 to 4: the same question, asked by accounts holding different grants, returning different evidence. Build the vault first — scripts/skykeep.sh demo takes a clean clone to this page in one command — and read what the demonstration deliberately weakens before showing it to anyone.
1. Reach is a number, and it is smaller than the vault
- Sign in as user_hr
- Document stats (the chart icon in the left rail)
- Count the documents: 184, on the packaged corpus
The packaged corpus holds 269 documents. This account can see 184 of them. Check the note at the top of this page first: if this deployment installs a corpus of its own, both numbers are about documents it does not hold, and you want your own two numbers here rather than these. Say both numbers out loud before asking anything — every later beat is a consequence of this one. Then say where the 184 came from: the corpus's twelve categories are split three ways — four categories (102 documents) in the hr compartment, four (85 documents) in sales, four (82 documents) in hr AND sales — and this account holds hr: 102 plus the shared 82. Nothing was hidden by the interface; the other 85 documents were never in this account's reach.
2. The spine: the same question, different grants
- Query (the magnifier icon)
- Ask the Civil War question from the sample questions below
- Repeat, signed in as user_sales, then as user_both
- Repeat once more as user_hr_ro — same answer, no upload
This is the demonstration. One question, asked three times, word for word the same, by accounts that differ only in their compartment grants. user_hr answers with the words spoken — the Proclamation, the second inaugural, the Gettysburg dedication. user_sales answers with the ground it was fought on — the Manassas battlefield and Arlington. user_both joins the words to the ground and cites documents from both compartments in one answer. No account was told what it was missing, because being told what you are missing is itself a disclosure. Ask it again as user_hr_ro, user_sales_ro or user_both_ro and the answer is word for word what the writing twin got: those accounts differ in what they may UPLOAD, not in what they may read.
3. The filter runs before the model, and you can watch it
- Query, and UNTICK 'All-In-One — retrieve and answer in one step'
- Step 1 — retrieval: press 'Retrieve chunks'
- Open 'Show the chunks this answer was built from'
- Step 2 — synthesis: press 'Synthesize over these chunks'
Two-step mode splits the seam open. Step 1 shows you every passage the model will reason over, and nothing outside your grants reaches it. Say it that way rather than 'the model sees exactly these words': each passage is widened to its enclosing section when the prompt is built, so the model may read more of a document than the box on screen shows — the same document, inside the same version, under the same permissions, but more of it. An audience that sees the answer quote a sentence which is not on screen should hear that explained, not discovered. Show the two per-compartment tallies next to the count: every compartment named there is one this caller belongs to. The mode is one fewer click, never one fewer check: step 2 receives chunk REFERENCES and the server re-resolves each one under the caller's own predicates on both paths.
4. Ask for something this account cannot reach
- Still as user_hr, ask the bats-of-Mexico question below
- Read the answer, then ask about a subject in no document at all
The two answers have the same shape. An account reaching past its compartments gets what an account asking about a document that was never written gets: nothing found. There is no 'you are not allowed to see this' anywhere in the product, because that sentence would confirm the document exists — which is the fact being protected. This is the beat that most audiences do not expect, so let the silence sit for a moment before explaining it.
5. The vault recorded you doing all of that
- Audit (open to every signed-in account only when the demonstration env file sets the audit-view phrase)
- Find the last five minutes
Reads are audited, not just writes — the queries just run, the documents they touched, and the refusal from beat 4, which is recorded as a denial rather than as an absence. Point at the presenter's own rows: the person demonstrating the vault is in the trail like anybody else. Then say what normally guards this screen — it is administrator-only, and it is open here only because the demonstration env file carries a second, separate acknowledgment phrase that no other setting implies.
6. The scan gates cannot be turned off
- Upload documents (the upload icon)
- Upload a file carrying prompt-injection-shaped text
- Sign in as admin001 and open quarantine review in /admin
Both ingest gates run on every upload, on every deployment, and NO configuration disables them — not a demonstration phrase, not an administrator, not an environment variable. A held document goes to quarantine for a human to look at; it is never silently passed and never silently dropped. An unconfigured scanner quarantines rather than waves through, which is the same fail-closed rule in its unhappiest form.
7. The model is a choice, and the console says what it costs
- Sign in as admin001 (on a deployment carrying SKYKEEP_DEMO_ADMIN_MFA_OFF this is the password alone; without it this account needs a mail relay)
- /admin, then the Settings → Models panel
- Read 'The two roles', then 'What this deployment has', then the consequences table
Two model roles, not one (ADR-0031): the answering model runs while somebody watches a spinner, the ingestion model runs at upload with nobody waiting, and they are chosen independently — a blank ingestion model resolves to the answering one, and the panel says which rather than leaving it to be guessed. Three things are worth pointing at. First, the list holds only what this deployment's runtime actually has, because a model that is not pulled cannot answer. Second, every claim is labelled curated, derived or measured, so an audience can tell a characterisation from an arithmetic consequence — the vault does not invent a description for a model nobody here has run. Third, the consequences table reports 'unknown' as a first-class outcome rather than dropping the row, because a missing row reads as 'nothing to worry about here'. Embedding is a deliberately separate control: changing the answering model changes the next answer, changing the embedding model invalidates every vector in the vault.
8. Close on what the demonstration did NOT weaken
- No clicks — this is the closing sentence
Everything weakened here widens who may sign in, or who may read the record. Nothing weakened here changes what the vault hands over. The compartment predicates are untouched, Postgres row-level security re-checks the identical predicate underneath the application, both scan gates run, and every access — allowed and refused — is written to a physically separate audit database. The packaged corpus's 184/167/269 reach numbers the demonstration opened with are produced by the same code a real engagement runs.
Sample multi-document questions
Questions that cannot be answered from one document, chosen so that the evidence falls on both sides of the compartment boundary. The corpus is split by CATEGORY: Biochemistry papers (hr), Candide essays (ODT) (hr), Candide essays (RTF) (hr) and HR policies (hr); Handwritten notes (sales), PV price graphs (sales), Product decks (sales) and Python PEPs (sales); Reflections on Greek myth (French) (both), Sales figures (CSV) (both), Spreadsheets (both) and User manuals (both). So user_hr reaches 184 of the 269 documents, user_sales 167, and user_both all 269 — and a question whose halves land in an hr category and a sales one returns materially different answers to the three accounts.
The table under each question names the three WRITING accounts, and it is complete: user_hr_ro, user_sales_ro and user_both_ro hold the same compartments as their twins, so each one gets the answer written on its twin's row. That is the point of them — the read-only accounts prove that what the vault hands over is decided by the compartment grants alone, while the level decides only whether the page offers an upload control.
Every question below names the documents its evidence falls in and the compartment each one lands in, and each was checked against the packaged corpus by converting the document with the product's own parser and finding the quoted phrase in the result. A title on its own was never taken as proof: a title is a claim about a document, and the converted text is the document as this vault actually reads it.
Ask your own as well, using the split above: pick a subject that appears in one hr category and one sales category, and ask it as user_hr, user_sales and user_both in turn. The questions below are worked examples of that move, not a script you are confined to.
How is a hiring decision supposed to be made here — and what has been sold to us to make it that way?
Documents it needs: Recruitment and Selection — Ninetail Studio (hr) for the rule the company set itself; Tinsley Sift — Tinsley People Tools (sales) for the product that sells the same discipline back to it.
| user_hr | The rule only: a vacancy authorised before it is advertised, the salary range published, applications sifted by at least two people against written criteria, and every candidate for a role asked the same core questions. This account cannot name one tool that would enforce any of it. |
|---|---|
| user_sales | The pitch only: a scorecard built from the role's requirements, the same questions handed to every interviewer, panel scores hidden until all of them are in. This account can describe the product and cannot quote the policy it would be sold against. |
| user_both | Both halves in one answer, citing a policy from one compartment and a sales deck from the other — the standard a company set itself, beside the software that claims to meet it. This is the clearest beat in the demonstration: one question, three answers, and the only thing that differs is the grants. |
Evidence: VERIFIED by exact phrase, each document converted with the vault's own parser. 'Every candidate for the same role is asked the same core questions and is scored against the same criteria' occurs only in Recruitment and Selection — Ninetail Studio (hr). "Sift builds a scorecard from the role's requirements and gives every interviewer the same questions" occurs only in Tinsley Sift — Tinsley People Tools (sales). No single compartment holds both halves.
Where are people expected to work, and how much warning does somebody get about next week's shifts?
Documents it needs: Remote and Hybrid Working — Meridian Fenwick Biosciences (hr) for the attendance rule; Quillfeather Rota — Quillfeather Software (sales) for the notice period.
| user_hr | The attendance rule only: three working patterns, hybrid roles on site at least two days a week, and the days fixed in a team's working agreement rather than chosen weekly. Nothing at all about rotas or notice — the scheduling material is out of reach. |
|---|---|
| user_sales | The scheduling half only: a rota product that holds the working-time rules as data, refuses to publish a schedule that breaks one, and treats the publication notice as a setting rather than a promise. It cannot say what any employer's attendance policy requires. |
| user_both | One answer covering both: the policy that says when people are on site, and the tool that says how far ahead they are told. A good second question precisely because neither half looks incomplete on its own — the audience has to be shown what was missing. |
Evidence: VERIFIED by exact phrase over the packaged corpus. 'Hybrid roles attend a location for at least two days a week' occurs only in Remote and Hybrid Working — Meridian Fenwick Biosciences (hr). "Two weeks' notice is a setting, not a promise" occurs only in Quillfeather Rota — Quillfeather Software (sales).
If information leaks, how fast must it be reported — and how would anyone find out we were exposed in the first place?
Documents it needs: Information Security and Data Handling — Calderwood Mutual Credit Union (hr) for the reporting duty; Whitlock Perimeter — Whitlock Assurance (sales) for the discovery side.
| user_hr | The internal duty only: four classes of information, a document taking the classification of the most sensitive thing in it, and a reporting clock measured in hours rather than days. This account cannot say how an exposure would be noticed. |
|---|---|
| user_sales | The discovery side only: continuous discovery of what an organisation exposes to the internet, findings ranked by what is actually reachable, and a weekly list short enough to finish. Not one word about what to do once something has leaked. |
| user_both | The whole loop — notice it, then report it — cited across the compartment boundary. Worth asking third, because by now the audience is predicting which half each account will be missing, and being right is what makes the point stick. |
Evidence: VERIFIED by exact phrase over the packaged corpus. 'a suspected incident reported in an hour is worth more than a confirmed one reported in a week' occurs only in Information Security and Data Handling — Calderwood Mutual Credit Union (hr). 'Perimeter discovers the estate continuously and ranks findings by what is actually reachable' occurs only in Whitlock Perimeter — Whitlock Assurance (sales).
What do these documents say about style — how a piece of writing should sound, and how code should be laid out?
Documents it needs: The Voice That Never Raises Itself (Amara Diallo) (hr) for the prose; PEP 8 — Style Guide for Python Code (sales) for the code.
| user_hr | The prose half only: a student essay arguing that the even, unruffled narrating voice of a satirical novel is its sharpest instrument. An argument about tone, and no rule about layout anywhere in reach. |
|---|---|
| user_sales | The code half only: a style guide for a programming language, up to and including its warning that a foolish consistency is the hobgoblin of little minds. Rules about indentation and line length, and nothing about how writing sounds. |
| user_both | Both, and the join is worth making rather than decorative: two documents from unrelated worlds that each argue style is a decision somebody made. Useful late in a demonstration, when an audience has begun to suspect the earlier answers were arranged in advance. |
Evidence: VERIFIED by exact phrase over the packaged corpus. "That flat narrating voice is the novel's most sophisticated device" occurs only in The Voice That Never Raises Itself (Amara Diallo) (hr). 'A Foolish Consistency is the Hobgoblin of Little Minds' occurs only in PEP 8 — Style Guide for Python Code (sales).
In the study of an inherited porphyria in a Spanish population, how many carriers actually developed the disease?
Documents it needs: High penetrance of acute intermittent porphyria in a Spanish founder mutation population and CYP2D6 genotype as a susceptibility factor (hr) — one document, and the number is printed in its results.
| user_hr | Answers it: penetrance of 52 per cent, with prevalence estimated at 17.7 cases per million inhabitants. |
|---|---|
| user_sales | Returns nothing. The research papers are not in this account's reach, and the honest answer is a decline. Worth showing on purpose: a confident wrong number here would look exactly like a right one to most of the room. |
| user_both | The same answer user_hr gets. |
Evidence: VERIFIED by exact phrase over the packaged corpus: 'Results: AIP penetrance was 52%, and prevalence was estimated as 17.7 cases/million inhabitants' occurs only in High penetrance of acute intermittent porphyria in a Spanish founder mutation population and CYP2D6 genotype as a susceptibility factor (hr).
One essay argues that a famous closing line is not advice to withdraw from the world. What is its case?
Documents it needs: What It Means to Cultivate Your Garden (Devon Okafor) (hr) — a single essay, and the whole argument is in it.
| user_hr | Answers it: the line arrives only after the characters have run out of both money and story, and the essay reads it as an instruction to stop explaining the world and start doing something in it — with work named as the guard against boredom, vice and need. |
|---|---|
| user_sales | Returns nothing; the essays are not in this account's reach. |
| user_both | The same answer user_hr gets. |
Evidence: VERIFIED by exact phrase over the packaged corpus: 'Martin says that work keeps away three great evils: boredom, vice and need' occurs only in What It Means to Cultivate Your Garden (Devon Okafor) (hr).
Is there anything in here that reads like a statement of taste rather than a specification?
Documents it needs: PEP 20 — The Zen of Python (sales) — one short document, and it is nothing but aphorisms.
| user_hr | Returns nothing: the language standards sit in the other compartment. |
|---|---|
| user_sales | Answers it, quoting a list of aphorisms that opens 'Beautiful is better than ugly'. A pleasant question to ask after a heavy one, and it shows retrieval landing on a very short document. |
| user_both | The same answer user_sales gets. |
Evidence: VERIFIED by exact phrase over the packaged corpus: 'Beautiful is better than ugly' occurs only in PEP 20 — The Zen of Python (sales).
What did a solar panel cost in 1975, and what does it cost now?
Documents it needs: Solar PV module price, 1975 to 2024 (logarithmic scale) (sales) — and this is the one question on the page whose honest answer is that the vault cannot read the document.
| user_hr | Returns nothing at all: the price graphs are not in this account's reach, so this is an ordinary refusal. |
|---|---|
| user_sales | Finds the document and cannot answer from it. Say it out loud in these words: the graph is an image; the vault stores it and cannot read its numbers. The file is sealed, listed and handed back on request, but no text is extracted from a picture, so there is no figure to quote. Show this beat deliberately — a product that invented a price here would be far worse than one that declines. |
| user_both | The same as user_sales: the file, and none of the figures printed on it. |
Evidence: VERIFIED by running the vault's own parser over the file: it ACCEPTS the image and stores it without conversion, reporting 'No text to extract without OCR (deferred).' There is therefore no converted text and no phrase to quote, which is exactly the claim being demonstrated — Solar PV module price, 1975 to 2024 (logarithmic scale) (sales) is retrievable as a file and unsearchable by content. Twenty-five documents in this corpus behave this way; all of them are graphs.
The coffee roaster's sales file — which regions does it break sales down by, over what period, and where is the money?
Documents it needs: Harbourline Coffee Roasters — regional sales by product (both) — one spreadsheet-shaped file in the shared third of the corpus.
| user_hr | Answers it: four regions — City, Suburban, Transport hubs and Wholesale — six product lines, and weekly rows from 2025-01-09 to 2025-03-20. Wholesale is the largest region by revenue and House Blend 250g the largest line. This file is in the shared third, so this account gets the whole thing. |
|---|---|
| user_sales | The same answer, and for the same reason. |
| user_both | The same answer again. |
Evidence: VERIFIED by exact phrase over the packaged corpus: the row 'Wholesale | House Blend 250g | HLC-0401-CA' occurs only in Harbourline Coffee Roasters — regional sales by product (both), and the four region names and six product names are the only ones in the file. NOT verified, and do not claim it on stage: the two ranking claims come from adding up all 252 rows, which the reader can do with the file in hand but which no language model answering from a handful of retrieved rows should be trusted to do. Ask for the regions and the dates; treat any total it volunteers as something to check against the file.
The cold-store workbook — what does it actually measure, and over how long?
Documents it needs: Cold Store Temperature Excursions — Grimsby Cold Chain Systems (both) — one workbook with two sheets.
| user_hr | Answers it: twenty-six chambers, measured two ways — mean chamber temperature in degrees Celsius, and minutes spent above the set point — week by week from 2024-W01 to 2025-W48. In the shared third, so every account gets it. |
|---|---|
| user_sales | The same answer. |
| user_both | The same answer again. Useful for showing that the vault reads a workbook as text, sheet by sheet, rather than treating it as an opaque attachment. |
Evidence: VERIFIED by exact phrase over the packaged corpus: the two sheet headers 'Chamber / Mean chamber temperature (degrees Celsius)' occurs only in Cold Store Temperature Excursions — Grimsby Cold Chain Systems (both). The second sheet header, 'Chamber / Minutes above set point (minutes)', occurs only in Cold Store Temperature Excursions — Grimsby Cold Chain Systems (both) too. The row labels run Chamber 01 to Chamber 26 on both sheets.
How do you make a search tool walk a whole directory tree, and what is the difference between the two ways of doing it?
Documents it needs: GNU Grep: print lines that match patterns (both) — a manual in the shared third, and the answer is in its options chapter.
| user_hr | Answers it: the lower-case recursive option processes every file under each directory operand but skips symbolic links met on the way down, while the upper-case one follows them; with no file operand at all the tool searches the working directory. |
|---|---|
| user_sales | The same answer. |
| user_both | The same answer again — and a good question for showing citations, because the evidence is a numbered option in a long manual rather than a paragraph of prose. |
Evidence: VERIFIED by exact phrase over the packaged corpus: 'How do I search directories recursively?' occurs only in GNU Grep: print lines that match patterns (both). The option name '--dereference-recursive' occurs only in GNU Grep: print lines that match patterns (both) as well.
What do the documents advise about typesetting a table — the rules drawn across it, and cells that span more than one row?
Documents it needs: The booktabs package: publication-quality tables in LaTeX (both) for the rules, and The multirow, bigstrut and bigdelim packages (both) for the spanned cells. BOTH compartments hold every document this question needs, which is the whole point of asking it.
| user_hr | The full answer. |
|---|---|
| user_sales | The full answer, and the same one. |
| user_both | The full answer, and the same one again. |
Evidence: VERIFIED by exact phrase over the packaged corpus: '1. Never, ever use vertical rules.' occurs only in The booktabs package: publication-quality tables in LaTeX (both). The phrase 'which provides a construction for table cells that span more than one' occurs only in The multirow, bigstrut and bigdelim packages (both). THIS IS THE CONTROL QUESTION and it belongs in the script. Without it an audience can reasonably conclude that user_hr is simply a degraded account. Run it and all three answers match — and so does each read-only twin, which reaches what its writing account reaches — which shows that the difference everywhere else is about GRANTS and not about some accounts getting a worse vault.
What do the documents say about the Apollo 11 moon landing?
Documents it needs: No document. Nothing in this vault answers it — not the engineering manuals (both), not the research papers (hr), not the sales material (sales). That is the point of asking it.
| user_hr | Nothing found, and the right answer is a plain refusal: the vault says it has no documents on the subject and cites none. |
|---|---|
| user_sales | The same refusal, in the same words. |
| user_both | The same refusal again — and this is what makes it worth asking. The widest account in the demonstration says no too. An account that declines because a subject is out of its reach and an account that declines because the subject is simply absent look identical from outside, which is what stops a refusal from being a hint about what else exists. |
Evidence: VERIFIED by exact phrase over the packaged corpus: the string 'Apollo 11' does not occur in the converted text of any of the 269 documents, the spaceflight engineering manuals included — those are about simulation tools and shell software, not missions. The phrase checked is the MISSION's name rather than the bare word, because two of the French essays name the god Apollon, and a bare 'Apollo' is a substring of that. This is the out-of-corpus check rather than the shared-third control above. A correct answer looks like 'I could not find anything about that in the documents you can reach', with an empty evidence panel beside it. A wrong answer is any sentence carrying a date, a name or a fact about the mission; if you see one, stop and open the evidence panel in front of the room, because making that visible is the entire argument for this product.
How these were checked, and what was not checked. Each question's evidence line reports exact-phrase presence in the packaged corpus: a committed archive of 269 fixed documents, split by category deterministically, and the phrase was looked for in the text this product's own parser produces from the document rather than in the original file. So 'this phrase occurs only in that document' is a property of the corpus and not a guess. That corpus is the one SkyKeep SHIPS, and a deployment can replace it — if the note at the top of this page says yours does, then these questions were never checked against anything you have installed, and the evidence lines below describe documents this vault does not hold. The compartment-boundary claims follow from that plus the category split stated above. Twenty-four of the 269 documents arrive held — uploaded and deliberately left unprocessed — and a held document is not retrievable, which is why no question above cites one. What is NOT guaranteed is the RETRIEVAL step: whether a given natural-language question surfaces the paper carrying its evidence depends on the embedding model, the chunk window and the deployment's own index, and it is not pinned by any test. If a question comes back thin, open the two-step mode and look at which chunks were admitted before blaming the answer. On a DEFAULT install the chunk window FITS the embedding model's input by construction — 190 words is 256 tokens at 1.3 tokens a word, which is what the shipped embedder accepts — so that is not the thing to suspect. It becomes worth suspecting only if an administrator has raised the window past the embedder's limit, in which case a chunk is indexed largely by its opening; the console says so on Settings then Models rather than leaving you to infer it.
Sample questions for the mixed-format corpus
A second corpus exists for deployments that want to exercise every format the vault reads rather than plain text alone: a fixed public-domain essay collection rendered one per document across .txt, .md, .html, .htm, .docx, .pdf, .pptx, .eml, .mbox, .tsv, .rtf, .epub, .odt, .ods, .odp, .xls, .ipynb, .rst, .json, .toml, .yaml, .yml, .xml, .ics, .srt and .vtt, plus a small generated records layer — three custody registers in .csv and a custodian directory in .xlsx. Its source and cache directory are configuration (SKYKEEP_DEMO_MIXED_CORPUS_URL, SKYKEEP_DEMO_MIXED_CORPUS_DIR).
Configured is not the same as built. This deployment holds the three settings demo_setup --mixed needs, so the corpus these questions draw on CAN be loaded here — but only a vault built WITH that flag actually holds it, and this page cannot see which way the vault in front of you was built. Ask one of these as user_both and check it returns something before you rely on any of them.
What makes them joins is that the records layer was built to hold the facts the essays do not. A document states its own reference code and matter code and never states who received it or when; a register states the custodian and the date and never states the argument; the directory is the only document anywhere in the corpus carrying an address, a desk number or an office, and it sits in sales alone. There is no single document that shortcuts any question below.
Who took custody of the paper arguing that the cure for faction is to extend the sphere, and how would I reach them?
Documents it needs: sk-010.txt (hr) for the argument and its reference code; register-hr.csv (hr) for the custodian and the date; custodian-directory.xlsx (sales) for the contact details. Three documents, three formats, two compartments.
| user_hr | The argument, the reference SK-010, the custodian's name (Alina Vestergaard) and the date it was received (2026-01-14) — and then it stops. The directory holding her address, desk and office is not in this account's reach, so it can name the person and cannot reach her. |
|---|---|
| user_sales | Almost nothing, which is the surprising half. It HOLDS the directory, so it can describe Alina Vestergaard's role and office perfectly well — but it holds neither the essay nor the hr register, so it cannot connect her to the question at all. Reach is not seniority; it is which documents. |
| user_both | The whole chain, in one answer that cites three documents in three different formats: the essay, the register row that dates it, and the directory row that says where to find her. |
Evidence: VERIFIED against the BUILT corpus, by exact phrase over the parsed text of every document — that is, after the vault's own parsers have read the .txt, the .csv and the .xlsx back. 'extend the sphere' occurs in sk-010.txt and nowhere else; 'SK-010' occurs only in sk-010.txt and register-hr.csv; 'Alina Vestergaard' only in register-hr.csv and custodian-directory.xlsx; her address only in custodian-directory.xlsx.
Which custodian took documents in BOTH matters, and how many in each?
Documents it needs: register-hr.csv (hr) and register-sales.csv (sales). Two documents, one in each compartment, and no third document that would let either be skipped.
| user_hr | Its own three custodians and their counts, and no way to answer the question asked. Worth saying out loud on stage: it cannot even tell that the question HAS an answer. |
|---|---|
| user_sales | The mirror image, with the other matter's three names. |
| user_both | Marta Kowalczyk — nine documents in M-2026-0141 and nine in M-2026-0207. The cleanest join in the corpus, because the answer is an intersection and an intersection taken against nothing is empty. |
Evidence: VERIFIED against the corpus builder: the hr roster and the sales roster share exactly one name, asserted directly rather than read off a comment, and the third slice's roster shares none with either. The counts follow from the assignment being a pure function of the document number.
Two papers here put a number on things: one says an army of no more than twenty-five or thirty thousand men, the other says the cure for faction is to extend the sphere. Give each one's reference code, custodian and received date.
Documents it needs: sk-046.txt with register-sales.csv (sales) for the first; sk-010.txt with register-hr.csv (hr) for the second. Four documents, both compartments.
| user_hr | The faction half only: SK-010, Alina Vestergaard, 2026-01-14. The army arithmetic is not in reach, so the answer is half an answer and says so. |
|---|---|
| user_sales | The army half only: SK-046, Marta Kowalczyk, 2026-02-19. |
| user_both | Both halves, with four documents cited. Run it straight after the previous question: the same two accounts that each held half of an intersection now each hold half of a list, and the shape of the shortfall is different. |
Evidence: VERIFIED by exact phrase over the parsed corpus. 'twenty-five or thirty thousand' occurs only in sk-046.txt and 'extend the sphere' only in sk-010.txt; every reference code reaches exactly two documents, its own and its register.
The hr register lists SK-004 as a Word document and SK-005 as a PDF. What does each say about foreign force, and do the two arguments agree?
Documents it needs: register-hr.csv (hr) for the two format claims, then sk-004.docx and sk-005.pdf (hr) for the arguments. Three documents in three formats — this is the question that shows the vault reads more than plain text.
| user_hr | The full answer, quoting a Word document and a PDF in the same breath. Show the citations: the point is that the evidence came out of a .docx and a .pdf, not out of a pre-flattened text file. |
|---|---|
| user_sales | Nothing. None of the three documents is in reach, and the refusal has the same shape as a question about a subject in no document at all. |
| user_both | The same answer user_hr gets. Deliberately: not every question has to divide, and one that does not, but still needs three documents and three formats, is the one to ask when somebody asks whether the vault can actually read a PDF. |
Evidence: VERIFIED by exact phrase over the parsed corpus, which for these two documents means after python-docx and pypdf have read them back. 'foreign force' occurs in sk-002.md, sk-003.html, sk-004.docx and sk-005.pdf — all hr — and in no other document in the corpus. Every document in the corpus is asserted to parse into non-empty text, so 'the PDF is readable' is a tested claim rather than a hope.
What does SK-073 argue, who took custody of it, and when?
Documents it needs: sk-073.txt and register-shared.csv — two documents, and BOTH of them are in the shared slice, so both are in every account's reach.
| user_hr | The full answer — SK-073, Noor El-Amin, 2026-03-18 — on the President's qualified negative, the veto. |
|---|---|
| user_sales | The full answer, and the same one. |
| user_both | The full answer, and the same one again. |
Evidence: VERIFIED by exact phrase: 'veto' occurs only in sk-073.txt, and SK-073's register row only in register-shared.csv. Both are in the shared slice and therefore in both compartments. THIS IS THE CONTROL QUESTION for this corpus and it belongs in the script for the same reason the plain-text corpus has one: it still needs two documents, so it is not a softball, and every account answers it identically, which is what proves the difference everywhere else is about grants rather than about some accounts getting a worse vault.
Denial pair — ask each account the half it cannot reach: 'What is Teodor Halvorsen's desk number and which office is he in?' and 'Which documents did Ruben Achebe take custody of, and on what dates?'
Documents it needs: custodian-directory.xlsx (sales) for the first; register-hr.csv (hr) for the second.
| user_hr | Answers the second from its own register — nine documents, each with its date — and returns nothing at all for the first. |
|---|---|
| user_sales | Answers the first from the directory — Teodor Halvorsen, +1-555-0193, Room 3F — and returns nothing for the second. This is the half worth dwelling on. It can see that Ruben Achebe EXISTS, because the directory lists him, and it still cannot say he handled a single document. Partial reach is not the same as no reach, and the vault does not paper over the difference. |
| user_both | Answers both, and can say the two people work in one team. |
Evidence: VERIFIED by exact phrase over the parsed corpus. Every custodian's address, desk number and office occur ONLY in custodian-directory.xlsx, which is asserted for all of them rather than spot-checked for one. 'Ruben Achebe' occurs in register-hr.csv (hr) and in custodian-directory.xlsx (sales) and nowhere else, so the sales side really does hold the person without holding his documents.
How these were checked, and what was not. Every evidence line above reports exact-phrase presence in the corpus AFTER the vault's own parsers have read each document back — so 'this phrase is in the PDF and in no other document' is a property of the built corpus rather than a guess about a generator. The compartment claims follow from the same slice boundaries the plain-text corpus uses (1-28 hr, 29-56 sales, 57-85 both) plus one deliberate placement: the custodian directory sits in `sales` alone, which is what makes the contact half of the first question unreachable for an hr-only account. What is NOT guaranteed is the same thing the plain-text questions do not guarantee — the RETRIEVAL step. Whether a natural-language question surfaces the document carrying its evidence depends on the embedding model, the chunk window and the deployment's own index, and no test pins it. And one thing more is not guaranteed here: this corpus is loaded only when demo_setup is asked for it BY NAME, with --mixed, from SKYKEEP_DEMO_MIXED_CORPUS_URL. A vault built WITHOUT that flag holds the plain-text corpus alone and will not answer these questions at all — the questions higher up this page are the ones for that vault.
Before anything else: what follows is HISTORY, not a current result. It was measured against a vault built by hand, at a time when no command this product shipped could install this corpus at all. demo_setup --mixed closed that afterwards, and these six questions have NOT been re-measured against a vault built by it — so read the rest as the last thing anybody actually observed rather than as what happens today. Measured, and it did not work. On 2026-08-26 all 89 documents were loaded into a fresh vault and every question above was asked as each of the three accounts. NONE of the six was answered — the vault retrieved essays and said, correctly, that the passages did not answer the question. Do not demonstrate with these questions until that changes. The corpus is not the problem and neither is the placement: asked in a register's own terms the register comes straight back. What these questions need is TWO hops — the essay gives you a reference code, and the code is what finds the custody row — and a single retrieval over the question's own words returns essays, because essays are what the question is made of. That is a retrieval design question, not a tuning one, and it is recorded rather than papered over. SUPERSEDED 2026-08-27: the second hop that sentence asks for was built, and these questions were still not answered. Re-measuring showed the earlier reading was wrong and that it was not a retrieval question at all — the vault under test held none of the documents the questions are about, because at that time nothing installed them — which is no longer so. The paragraph is kept because a wrong diagnosis that was honestly measured is worth more to the next reader than a deleted one.