Renderingen-US

Inferentia — local inference

Inferentia is a local-first inference server written in Faber. It is meant to load a GGUF language model from disk and serve it through a command-line tool and a small HTTP API, on your own machine, with no outside service.

It is an application, not a library you import. It is listed here because it is built on Gradus and Norma, which makes it a real test of both. It is in development and is not a released product; read the status section first.

Who owns what#

Inferentia owns the product: configuration, the command line and API contracts, the request lifecycle, streaming, and cancellation. The model itself is not here.

  • Gradus owns the machine-learning semantics: admitting a GGUF file, the tokenizer, and the generation loop. Inferentia declares it as a dependency in faber.toml.
  • Norma provides the HTTP server, headers, and server-sent events.
  • llama.cpp is a reference oracle for comparing output. It is not a dependency.

The design stance is to fail closed: an unsupported architecture, quantization, or tokenizer is an error with a typed cause, not a silent approximation.

Commands#

The command line has two subcommands.

inferentia serve --model <PATH> [options]
inferentia live  --model <PATH> [--prompt TEXT] [--ignore-eos]

serve binds an HTTP listener on the loopback interface. live loads a model, generates one completion for a prompt, prints the token ids and the decoded text, and exits.

OptionMeaning
-m, --model <PATH>Path to the GGUF file (required)
--host <HOST>Bind host; loopback only, default 127.0.0.1
--port <PORT>Bind port; default 8080, 0 picks a free port
--max-tokens <N>Default output token ceiling; default 32
--max-prompt <N>Prompt-token admission bound; default is the model's limit
--context <N>Session context ceiling; default is the model's limit
--seed <N>Deterministic seed; default 1
--log-level <LEVEL>info, warn, or error on stderr
--expected-digest <SHA256>serve: require the admitted model's digest to match

Usage errors are reported before the listener binds.

The HTTP API#

serve answers three paths:

PathWhat it does
/healthReports starting, then ready with the model identity, or failed with the typed cause
/modelThe admitted model's identity facts, or 503 before the server is ready
/generateTakes a JSON request, validates it against Gradus's admission rules, and returns JSON

The server binds first and admits the model afterward, so /health can report progress. /generate returns JSON by default. With Accept: text/event-stream or ?stream=1 it sends one server-sent event per token; generation itself is still a batch call, so this is a streaming format, not streaming compute.

Status#

As of 2026-09-30:

LayerState
Command line, configuration, and /health, /model, /generateWritten, in one source file of about 2,850 lines
Model admission, tokenizer, generationProvided by Gradus
Package checkDoes not pass today: Inferentia imports Gradus, which currently fails faber check
End-to-end resultThe one recorded live run did not match its expected output (see below)
HardwareThe current work targets the CPU path. GPU execution is planned separately
Speed and qualityNo claims. No measurements are published
Deployment and machine provisioningA separate work stream

The recorded run. On 2026-08-18 a live run loaded a 360-million-parameter SmolLM2 model (Q4_K_M quantization) and generated with greedy decoding and a fixed seed. The first token matched the expected output; the second was the end-of-sequence token, so generation stopped after two tokens instead of the expected sixteen. The verdict was recorded as a failure. The run took about 22 minutes in a debug build. The repository's evidence directory holds no later receipt.

Inferentia is a good place to read how a real Faber application is put together, and an honest picture of where local inference stands. It is not yet a tool to run models with.

Source: github.com/faberlang/inferentia.