Skip to content
NewEraAI

Tools

LFM2.5 DSpark update targets faster inference for AI services

Liquid AI released LFM2.5 DSpark, positioned for up to 3.2 times faster inference. Teams running LLM workloads can use this update to reduce latency in production.

20 August 2026

Close-up of a computer screen displaying ChatGPT interface in a dark setting.
Photograph by Matheus Bertelli · Pexels

A new inference focused update for LFM2.5 has been released, aiming to cut response times. If you run LLMs in production, the immediate question is how to measure the latency change in your own workload.

What changed

The release introduces LFM2.5 DSpark and claims up to 3.2 times faster inference. The stated goal is improved runtime performance during inference.

Why it matters for business teams

Lower inference latency can directly affect customer workflows that depend on fast responses. It can also reduce operational time for chat, support triage, and other interactive tasks where users wait on generation.

What to do next

Run a controlled benchmark on your traffic profile and measure end to end time, not just model compute time. Compare the current baseline to the new LFM2.5 DSpark setup, and track quality side by side so you know performance gains do not come with unacceptable changes.

Next step

Start with the free AI Opportunity Assessment.

A short, no-obligation conversation about where enquiries, hours and revenue leak today. You do not have to pick a tier to have it, and what comes out of it feeds Discover, so the first paid day starts from evidence rather than a blank sheet.