← Back

Model API

An OpenAI-compatible inference endpoint for open-weights models, served on hardware we operate. Same request and response shapes as the OpenAI Chat Completions API, including server-sent-event streaming and token usage accounting.

Models

Inkling Small LAUNCH

Thinking Machines Inkling-Small — 276B total parameters, 12B active, mixture-of-experts. Apache-2.0. Served at NVFP4 on 2× NVIDIA H200 with tensor parallelism, prefix caching enabled.

GLM-5.2 PLANNED

GLM-5.2, targeted at long-context agentic workloads. We have characterised it on our own multi-GPU hardware and hold a serving configuration that materially outperforms the obvious one for this class of machine. Capacity is being provisioned; timing will be announced here.

Further models EVALUATING

We add models where we can serve them well rather than broadly. Additions are announced here and reflected in the models endpoint above. We publish the quantization and the context length we actually serve — never the model's theoretical maximum.

How we operate

Access

The Model API is currently available through inference routing platforms. Our listing is pending review.

Self-serve API keys and credit top-up are coming soon. In the meantime, if you would like direct access, volume pricing, or an evaluation key, email contact@neuralaccel.com and we will set you up manually.

Policies

Use of the Model API is governed by our Privacy Policy and Terms of Service.