pharo-infer

PharoInfer

Pharo 13 & 14 License: MIT PRs Welcome Status: Active

PharoInfer is a fully in-image inference engine for Pharo Smalltalk. It loads a GGUF model file directly from disk and drives llama.cpp through UFFI — there is no HTTP server, no Ollama bridge, and no subprocess. Talk to the model straight from the image.

Requirements

model.gguf is the model weights only. The client machine also needs the native llama.cpp runtime, either installed by the user or shipped with your Pharo image/application.

Installation

Load the package in Pharo. If you are using this local checkout:

Metacello new
  githubUser: 'pharo-llm' project: 'pharo-infer' commitish: 'main' path: 'src';
  baseline: 'AIPharoInfer';
  load.

Run a GGUF model from Pharo

Do these steps in order. The .gguf file alone is not enough; the client machine must also have the native llama.cpp runtime built or bundled.

1. Build the native runtime

In a terminal, from this repository:

cd /path/to/pharo-infer
sh scripts/build-native.sh

The script clones/builds llama.cpp as a shared library and compiles the small PharoInfer shim into $HOME/pharo-infer-native/lib. It supports macOS, Linux, and Windows. On Windows, run it from Git Bash or another POSIX-compatible shell with CMake and a C/C++ compiler available.

On macOS this creates:

$HOME/pharo-infer-native/lib/libai_llama.dylib

On Linux this creates:

$HOME/pharo-infer-native/lib/libai_llama.so

On Windows this creates:

$HOME/pharo-infer-native/lib/ai_llama.dll

Keep the whole $HOME/pharo-infer-native/lib folder together. It contains libai_llama plus the llama.cpp / ggml libraries it depends on.

2. Put a model on disk

Use a GGUF model file:

mkdir -p "$HOME/pharo-models"
cp /path/to/model.gguf "$HOME/pharo-models/model.gguf"

3. Point Pharo at the native library

Pharo will look for libai_llama.so, libai_llama.dylib, or ai_llama.dll on the default library search path, and also under FileLocator imageDirectory / 'pharo-infer-native' / 'lib' and FileLocator home / 'pharo-infer-native' / 'lib'. To override, pin it from the image:

macOS:

AILlamaLibrary libraryPath:
  (FileLocator home / 'pharo-infer-native' / 'lib' / 'libai_llama.dylib') fullName.

Linux:

AILlamaLibrary libraryPath:
  (FileLocator home / 'pharo-infer-native' / 'lib' / 'libai_llama.so') fullName.

Windows:

AILlamaLibrary libraryPath:
  (FileLocator home / 'pharo-infer-native' / 'lib' / 'ai_llama.dll') fullName.

4. Load the model and ask it something

Run this in a Pharo Playground:

| backend manager engine model answer |

backend := AILocalBackend new
  nThreads: 8;
  contextSize: 2048;
  batchSize: 512;
  nGpuLayers: 0;
  yourself.

manager := AIModelManager new.
manager currentBackend: backend.

model := manager loadModel:
  (FileLocator home / 'pharo-models' / 'model.gguf') fullName.

engine := AIInferenceEngine new.
engine modelManager: manager.

answer := engine
  complete: 'Say hello from Pharo in one short sentence.'
  model: model name.

Transcript show: answer; cr.
answer

For Apple Silicon / Metal GPU offload, try this before loading the model:

backend nGpuLayers: 999.

More examples

Streaming

engine
    stream: 'Tell me a joke about Smalltalk'
    model: model name
    onToken: [ :piece | Transcript show: piece ].

Chat

| request |
request := AIChatCompletionRequest
    model: model name
    messages: {
        AIChatMessage system: 'You are a helpful AI assistant.'.
        AIChatMessage user: 'What is Smalltalk?' }.
AIChatAPI default complete: request.

GPU offload and threads

AILocalBackend new
    nGpuLayers: 999; "offload all layers"
    nThreads: 8;
    contextSize: 4096.

Architecture