Introducing Juno-N & Locai One

    James DraysonJames Drayson
    11 August 2026
    7 min read
    Introducing Juno-N & Locai One

    Today, we're launching Juno-N and Locai One.

    Think of Locai One as AI in a box: an on-prem AI appliance built on NVIDIA AI infrastructure and NVIDIA Nemotron open models. Locai One provides a practical and cost-effective route to own and customise enterprise AI by providing a private AI data centre in your office.

    Locai One combines the hardware, our Juno model series and Locai OS, our AI operating system, in one integrated appliance. Models and data can remain on infrastructure you control, while employees, applications and agents connect to it across your own network.

    Our mission at Locai Labs is to bring AI out of data centres and into your organisation, so you can own your own intelligence.

    Compress the model far enough and the data centre becomes the box in the corner of your office. That box is Locai One.

    Hardware, models and an AI operating system

    Locai One comes in two configurations: Locai One and Locai One Pro, with one or two NVIDIA Blackwell GPUs.

    It plugs into an organisation's existing network and provides local compute for AI workloads including coding, document analysis and agents.

    But the hardware is only one part of the product.

    Introducing Juno-N & Locai One
    Locai One and Locai One Pro

    Every Locai One comes with access to our Juno model series, models compressed and optimised specifically for local inference on the appliance. Organisations can also run compatible open-source models from Hugging Face, so Locai One is not locked to a single model family.

    The third layer is Locai OS, our AI operating system, which handles model serving, APIs, access, the Locai App and the inference layer that turns the underlying hardware into an AI appliance.

    Introducing Juno-N: built on NVIDIA Nemotron 3.5 Lightning

    Launching alongside Locai One is Juno-N-Coder-25B, our latest Juno model, built on NVIDIA Nemotron 3.5 Lightning.

    Locai Labs was one of NVIDIA’s early-access partners for Nemotron 3.5 Lightning, and we want to thank the NVIDIA team for giving us early access to the model and supporting our work.

    We were particularly excited by Lightning because of how well its architecture lends itself to what we are trying to achieve with Juno.

    Nemotron 3.5 Lightning punches well above its size on coding and agentic workloads while being lightning fast, exactly the capabilities that matter for long-running agents. Due to its fully open training data and sparse MoE architecture, it makes an excellent model for pruning and post-training. Using our SPACE compression algorithm, we compressed it to Juno-N-Coder-25B, quantised to NVFP4 for on-device inference, while retaining competitive coding and agentic behaviour.

    But the goal with Juno-N was not simply to make Nemotron smaller.

    The goal was to make it specialised.

    For Juno-N-Coder-25B, we specifically targeted two areas: software engineering and agentic workloads. SPACE was used to preserve the parts of the model responsible for those capabilities, while stripping back capability outside the domains the model was being built for.

    Introducing Juno-N & Locai One
    Juno-N-Coder-25B Results - Table
    Introducing Juno-N & Locai One
    Juno-N-Coder-25B Results - Spider diagram

    The result is visible in the model’s performance profile. Juno-N retains performance where we deliberately targeted it, particularly across Software Engineering and Agentic tasks, while its capability profile is reduced outside those target areas.

    That narrower profile is intentional.

    We are not trying to squeeze an entire general-purpose frontier model into a smaller box. We are identifying the capabilities an organisation actually needs, preserving those capabilities, and removing more of what it does not.

    How SPACE turns large general purpose models into smaller specialised ones

    All compression involves trade-offs.

    SPACE is designed to preserve the subnetworks in the model that correspond to the capability you actually need.
    Introducing Juno-N & Locai One
    SPACE: Locai Labs' model compression algorithm

    Juno-N is the clearest example of what that means.

    Rather than attempting to preserve Nemotron 3.5 Lightning equally across every benchmark and every domain, we targeted the capabilities required for coding and long-running agents. The resulting model remains competitive in those areas while becoming deliberately more specialised elsewhere.

    That is an important distinction.

    Traditional model compression asks: how much of this model can we remove while keeping the model as similar as possible?

    With SPACE, we are asking a different question:

    What does this model actually need to be good at?

    By preserving the subnetworks associated with those capabilities, we can turn large general-purpose models into smaller models designed for specific workloads and hardware.

    For Juno-N, that means software engineering and agents. Over time, our goal is to apply the same approach across different models and domains, creating a family of specialised Juno models that can run efficiently on local hardware.

    We also believe organisations should have choice over the models they run.

    Juno-N-Coder-25B is available openly on Hugging Face, and our Juno models will be open-sourced so organisations can inspect them, keep them and build on them.

    The Juno family will continue to expand across different base models and workloads, with Juno-D and Juno-I coming next.

    And customers are not restricted to Juno. Locai One can run compatible open-source models from Hugging Face, giving organisations the freedom to choose the model that best fits the job.

    This is ultimately what makes the Locai One possible. Instead of bringing ever-larger amounts of compute to the model, we are shrinking and specialising the model for the compute available inside the appliance.

    Locai OS: our AI operating system

    Running AI locally should not mean having to build your own data centre software stack.

    Locai OS is our Linux-based AI operating system. It handles the complexity of running, serving and connecting to local AI so Locai One behaves like an appliance rather than a data centre project.

    A Locai One can be set up in around 15 minutes. Plug it into a wall socket and your network, choose your model, configure your API access and connect it to your existing applications.

    Locai OS includes an admin portal, the Locai App and a Rust-based inference engine that turns the GPUs into a token factory, allowing Locai One to serve many users and workloads at once.

    OpenAI- and Anthropic-compatible APIs mean existing applications, coding tools and internal agents can connect to Locai One without organisations having to rebuild their integrations, and without model-provider token charges for every request.

    Introducing Juno-N & Locai One
    Locai OS, Locai Labs' AI operating system

    Access your Locai One from the Locai App

    The Locai App is how employees interact with their organisation's Locai One from their laptop or desktop.

    Employees can ask questions, write and review documents, analyse files and work with their organisation's AI through a familiar interface, while the model itself runs back on the Locai One.

    The appliance can sit securely inside the organisation while employees connect to it across the local network. Remote access can also be enabled by the organisation when required.

    Introducing Juno-N & Locai One
    Locai App

    Why bring AI in-house?

    Most organisations currently rent their AI from the cloud.

    Every prompt, document and agent action is processed on infrastructure owned by somebody else, using models the organisation does not control.

    Those models can change. They can be downgraded or discontinued. Pricing can change. And as AI becomes more deeply embedded into an organisation, that dependency becomes increasingly important.

    Locai One changes that relationship.

    Instead of paying a model provider for every token, an organisation owns a known amount of local computing capacity. As usage grows, another employee question, document or agent step does not create another model-provider inference bill.

    Proprietary documents, code and internal context can remain on infrastructure the organisation controls. The organisation decides which models run, who can access them and when those models are updated.

    The more important AI becomes to an organisation, the stronger the case becomes for owning the infrastructure and models that intelligence depends on.

    Built with the NVIDIA ecosystem

    Locai One is built on NVIDIA AI infrastructure, with NVIDIA Blackwell GPUs providing the compute underneath the appliance.

    Our first Juno model is built on NVIDIA Nemotron 3.5 Lightning, and we are excited to keep working with NVIDIA as we push larger AI models onto increasingly smaller amounts of hardware.

    Ecosystem support

    "The Incubator for AI is pleased to be trialling Locai One and to support the kind of homegrown AI innovation that companies like Locai Labs represent, it's a great example of Britain's AI sector at work"

    Max Hollingdale, Head of Applied AI Engineering, Incubator for AI, UK Government

    AI in a box

    Locai One is available to pre-order from 11 August 2026, with early customers able to secure one of our first production builds.

    Our goal is simple: bring AI out of the cloud and into your organisation, so you can own the intelligence you depend on.

    AI in a box. Built to run in-house.

    Enjoyed this article?

    Share it with your network