Image description

Large language models are fluent, useful — and they make things up, often with total confidence. If you're shipping anything on top of an LLM, you eventually hit the same question: can I tell when a response is likely wrong before a user sees it?


That question is why we built our first product at Komplex AI: a hallucination detector for LLM output.


What it does


You give it an AI-generated response, and it returns two things: a calibrated probability that the text is hallucinated (0 to 1), and the type of error — a fabricated fact, a fake or misattributed citation, a misleading-but-technically-true claim, a false refusal, and a few others.


So instead of a vague "this might be wrong," you get a number you can threshold on and a category you can route on — gate the answer, flag it for review, or trigger a retry.


Use it however you build


The detector is a hosted service you can reach in whatever way fits your stack: a web app (paste a response and get a score, no setup), a REST API (one call to /api/detect), SDKs for Python (pip install halu) and JavaScript/TypeScript (npm i @komplexai/halu), a LangChain integration (langchain-halu) to annotate, gate, or auto-retry generations inside an LCEL chain, an n8n node to drop a "Detect Hallucination" step into a no-code workflow, and an MCP server so Claude and other agents can call the detector as a tool.


The SDKs and integrations are open source (Apache-2.0); the detector itself is a hosted API. There's a free tier (no credit card), and we don't retain the text you submit.


What it is — and isn't


We think honesty about limits is part of the product, so we'll say it plainly. It's a probabilistic signal, not a fact-checker: it estimates risk, it doesn't look anything up. It works on English natural-language responses, up to about 2,000 characters, and is best on complete, multi-sentence answers — very short inputs can be over-flagged. The first call after a quiet period can take 10 to 30 seconds to warm up, then it's fast.


Per-error-type accuracy is published on the Performance page (detector.komplexai.io/performance) rather than hidden behind a single headline number.


Try it


It's free to try, and it takes about ten seconds: paste any AI response and see what comes back.


detector.komplexai.io