A Java and JVM port of llama.cpp using jextract, enabling local large language model (LLM) inference through native foreign function and memory API interop. Natively supports macOS M-series and Linux x86_64 with GPU acceleration. Platform and hardware support (Windows, ARM, CUDA, etc.) can be extended through custom builds.

Artifacts using Llamaj CPP (3)
Sort by:Popular▼

GLiClass decoder-kv (Qwen3) family: causal backbone on llama.cpp via llamaj.cpp, scorer on ONNX Runtime
Last Release on Sep 16, 2026
llama.cpp implementation of the Inference API
Last Release on Sep 24, 2026
llama.cpp implementation of the Inference API
Last Release on Jun 6, 2026
  • Prev
  • 1
  • Next