palOMine · Personal AI appliance

Your own AI. Your own hardware. It gets better, not just bigger.

palOMine is a personal AI appliance that runs entirely on hardware you own. Coding agent, voice, image generation, chat bots — one box, zero API keys, nothing leaves your network.

Most AI agents just pile up context until they rot. palOMine’s memory decays what you don’t use, flags contradictions instead of silently overwriting facts, and distills whole sessions into reusable skills — so it gets sharper over time instead of slower.

Every surface — CLI, chat, voice, Telegram, Discord, Signal — runs through the same bounded-authority core. You decide what it’s allowed to touch; remote requests can’t mutate anything by default.

Join the waitlist for the first production run.

palOMineEVO X-2 · 128GB
online · 128GB unified memory
agent · skills · memory
USB-C
USB-A
HDMI
RJ45
SD
PWR
  • Runs fully offline on your own GMKtec EVO X-2 (128GB) — no cloud, no subscription API
  • One agent, every surface: terminal, web, voice, and your chat apps
  • Ships pre-configured — plug in, connect your accounts, done
Single SKU

GMKtec EVO X-2, 128 GB.

One model, one configuration. Open-weight models run on a single preconfigured mini-PC — no proprietary stack, no build-to-order rack, no vendor lock-in.

Nemotron-3.5-Lightning
Qwen3.5-2B
Whisper-Large v3 Turbo
Kokoro
Moonshine-Medium
FLUX-2-Klein
RealESRGAN x4+
nomic-embed v1
bge-reranker v2 m3
See supported models
Model
GMKtec EVO X-2, 128GB unified memory — the single SKU
Form factor
Mini-PC desktop: sit it on a desk or shelf, plug in power + ethernet
GPU
Built-in integrated GPU running open-weight models — no separate GPU to configure
Quantization
Whatever the model ships with from upstream weights; no customer choices
Models
Reasoning/coding — NVIDIA Nemotron-3.5-Lightning 30B (GGUF). Plus Qwen3.5-2B (background), Whisper-Large v3 Turbo (STT), Kokoro (TTS), Moonshine-Medium Streaming (streaming STT), FLUX-2-Klein (image generation), RealESRGAN x4+ (upscaling), nomic-embed v1 (embeddings), bge-reranker v2 m3 (reranking). All pre-loaded; not a swap-in catalogue.
Network
Fully offline; nothing leaves the appliance unless you explicitly forward traffic
Monitoring
Local logs + status page on the box itself — no third-party agent or vendor sink
Data plane
Prompts, embeddings, and run state stay on-device; local storage encrypted at rest
Real-world performance

Sustained 58.14 tok/s on the appliance — not a burst number.

Measured on Nemotron 3.5 Lightning 30B-A3B MXFP4 running on the GMKtec EVO X-2 — not a synthetic peak, not a marketing number. The appliance holds steady across long turns instead of throttling mid-conversation.

Once a conversation crosses ~9–10k context tokens, the box reuses 98–99% of the cached context on the next turn — TTFT drops to ~0.5–0.6s instead of paying the multi-second prefill again. Long, persistent conversations stay fast turn to turn instead of slowing down, which is the differentiator versus typical local/on-prem AI setups where context reprocessing makes long exchanges progressively slower.

Workload model
Nemotron 3.5 Lightning 30B-A3B MXFP4
Sustained generation
58.14 tok/s average across long turns; holds mid-50s over a single 5,600+ token generation — not a short burst.
Prompt processing (prefill)
~1,204–1,260 tok/s on long uncached requests.
TTFT — cold, uncached
  1. 5.9k context4.72s
  2. 7.0k context5.62s
  3. 7.85k context6.52s
Context caching
Once context reaches ~9–10k context tokens, consecutive turns reuse 98–99% of the cached context — TTFT drops to ~0.5–0.6s.

Reserve a unit for the first production run.

One email and you’re on the waitlist. We notify signups ahead of the first hardware run.