Skip to content
SkyKeephelp

Choosing which models the vault runs

SkyKeep runs every AI step against a model runtime inside your own

deployment, and an administrator chooses which models. There are two

choices, not one, because the jobs have opposite shapes: reading a

document at upload happens once per document with nobody waiting, while

answering a question happens once per question with a person watching.

The vault will only offer models the runtime actually has, and it

labels every claim it makes about them — whether it characterised the

model itself, derived what it says from the size the runtime reported,

or measured it on this deployment. It will not invent a description of

a model nobody here has run.

Choosing a model chooses who writes the sentence. It never changes what

the sentence may be written from: your compartment grants decide that,

before any model is asked anything.

The ingestion model borrows the answering model, and says so

A vault told which model to read with uses that one

The embedding model never quietly borrows a chat model

A model the vault has characterised says what it is good and bad at

A model nobody here has run is described only by its own numbers

A model whose size the runtime never reported gets no description

A runtime that cannot be reached offers nothing rather than guessing

An embedding model chosen to answer questions is refused

A chunk longer than the embedding model can read is reported

A coupling the vault cannot judge says so rather than passing

A prompt is never silently truncated to fit the model

A prompt that fits is not marked as truncated