A GenUI dashboard went from 54 seconds to 3.8, at zero generation cost, with the same completion rate. Andrew King plugged TypeSafe AI's Jev into a plant-monitoring demo. Jev is a decision model: you hand it a state and typed questions, it returns typed answers with confidence scores, in parallel. It generates no text and never touches a measured value. What he caught early was the tell. The data was right and the reasoning was right, but the frontier model was making zero tool calls. The data was already fetched. The model was retyping numbers into a template and deciding what goes on screen, one token at a time. That's classification work. Jev answers about 100 closed questions in half a second. Code fetches the data, assembles from a pre-approved component catalog, and fills in every number at render. Most of the 3.8 seconds is the query. The limits are real: synthetic plant data, prose from templates, and anything outside the catalog falls back to an LLM. In a regulated environment that hard ceiling on what can render is a feature. Before swapping models, list every decision your composer makes. If more than half are closed questions with a fixed answer set, you're paying generation prices for classification. Full breakdown with benchmark table at the link in comments! #vgv #jev #genui
GenUI dashboard speeds up to 3.8 seconds with TypeSafe AI's Jev
More Relevant Posts
-
🧠 Jev on VegaDūta.ai — Decision Intelligence, Not Another Provider Does Jev act like another provider? No — and that’s deliberate. Providers (OpenAI, Anthropic…) generate the text your users read. Jev never writes a word. It answers typed questions — noul (yes/no probability), choice (pick an option), score (rate on a scale) — and only the platform’s internal decision layer consumes it. That’s why Jev isn’t in the provider list and can’t be misconfigured on an agent. 🔌 How VegaDūta Uses Jev 1️⃣ Cost Routing (QueryComplexityClassifier) User message → heuristic says “probably complex” Jev noul: “can a small, cheap model handle this well?” ≥0.75 yes → route to ECONOMY model instead of frontier. 💡 Real case: “What are your hours?” → heuristic would waste a frontier call. Jev says 0.9‑simple → runs on the cheap model. Thousands of turns = real money saved. 2️⃣ Quick Task Agent Selection (AgentRelevanceSelector) “Book a table at Mario’s Friday” → Jev choice over tenant’s agents. Picks reservations‑bot semantically, no keyword hacks. One batched noul call flags secondary‑relevant agents. 💡 Real case: 12 agents → old path misrouted on word overlap. Jev understands intent and routes correctly. 3️⃣ BYOK + Playground (/jev) Tenants bring their own TypeSafe key. Experiment with noul/choice/score directly. See raw JSON answers, model version, usage, and key source. Failures are visible, not hidden. ⚡ Why This Shape Ultra‑cheap: ~$0.042 per 1M input tokens, no output charge. Fast & safe: tight 2s/3s timeouts, always fail‑open → never blocks, only improves. Isolated: per‑tenant key + cache scope → no leaks across tenants. ✨ Out of the box, Jev is the System‑One layer in front of your expensive System‑Two models — fast, cheap decisions that save money and improve routing. 👉 Explore Jev now on https://vegaduta.ai #AgenticAI #DecisionIntelligence #OpenSource #BYOK #Automation #AIPlatform #Innovation #FutureOfAI #QuantumLeap #VegaDūta #Jev #AgenticAI #DecisionIntelligence #SystemOneAI #BYOK #AIPlatform #Automation #CostRouting #QuickTaskSelection #Innovation #FutureOfAI #OpenSourceAI #VegaDūta
To view or add a comment, sign in
-
Everyone's sharing benchmark screenshots showing open-weight models have "basically caught up" with closed frontier models. I read a deep dive this week that pushed back on that framing, and the argument holds up. The aggregate gap (per Epoch AI's capability index) is around 4 months behind the closed frontier, and it hasn't been shrinking. It grew from about 3 months in 2023-2025 to roughly 4 months by mid-2026, and the authors themselves warn that number probably understates the real gap, since open models get tuned against public benchmarks while closed labs keep their best systems private. But the more useful point isn't the number, it's that "benchmark parity," "parity on your specific task," and "parity in production" are three separate claims. A model can be indistinguishable from the frontier at extracting fields from a document and fall apart on long agentic chains with tool use, since that's exactly where closed labs are pouring their effort. One team in the article ran the same open model on the same week: fine on classifying disputed payment cases, then confidently citing internal rules that were never in the context when asked to generate explanations. Two things stood out to me as an engineer: - Pricing isn't a fixed variable. The article tracks one model's price over five months: preview price, a "permanent" 75% discount, a silent rebuild under the same model name, then peak/off-peak pricing that made the off-peak rate roughly 4x more expensive during exactly the hours a lot of non-US traffic runs. The endpoint name isn't a stable identifier for what's actually serving your requests, and the API sitting in front of open weights is still an external dependency with its own changelog. - "Open" describes a license, not a guarantee. Weights under MIT or Apache 2.0 are genuinely yours once downloaded. Weights under a vendor's custom license with revenue thresholds are a different commitment entirely, and that's before your security team even weighs in on where the data is allowed to go. None of this is an argument against open-weight models. We use them for classification and RAG retrieval where the numbers hold up cleanly. It's an argument against treating a benchmark screenshot as a migration decision. #AI #LLM #MachineLearning #Engineering
To view or add a comment, sign in
-
-
Came across TypeSafe AI's Jev (https://typesafe.ai/) this week, and I think the pattern behind it matters more than the model itself. The idea is to stop making decisions with a text generator. You send a state and a set of typed questions, and get back structured answers your code branches on directly. Choice picks an option. Score rates against a rubric. Noul answers true or false. Each one comes back with a probability distribution and a confidence value. Nothing is generated, so nothing needs parsing. Most of the examples around it lean real-time: smart home automation, game state, decisions at 150ms. Impressive, but I think it there are many more use cases where this is most useful.. Look at where the decisions actually sit in an ordinary RAG pipeline: User query ↓ route it — which index, which tool, which path (Choice) Vector search ↓ rerank — score each chunk before it enters the prompt (Score) The LLM writes the answer ↓ verify — is every claim supported by the retrieved context? (Noul) Ship it, or send low confidence to a human Three of those seven steps are decisions, not writing. Today most of us make all three by prompting an LLM to emit JSON and hoping it parses. The numbers are what make it practical: → $0.042 per million input tokens, output tokens free → around 150ms per call → questions evaluated in parallel against a single state ingest, so ten checks cost roughly what one costs → worked example: an 8k-token state with six grounding checks lands near $0.0004 per turn At that price the guardrail stops being a budget conversation. You can verify every answer instead of sampling five percent of them for review. I haven't run this in production yet, so I'm curious about the part the docs can't tell me. How does confidence calibrate on your own data? If you've put Jev on a real workload, I'd like to hear where it held up and where it didn't. #RAG #AIEngineering #LLMOps #GenAI #TypeSafeAI
To view or add a comment, sign in
-
Most curated tool lists collect adoptions. This one opens with the teams that measured the model and decided not to use it. That inversion is what makes awesome-jev worth a look. I recently came across awesome-jev, a catalog of 1,207 public resources for Jev, the System One decision model from TypeSafe AI. What caught my attention is not the size. It is the ordering. The section called "Measured, not claimed" leads with rejections. Nous Research's Hermes Agent ported Jev's compaction approach, measured it against the summariser they already ship, and published the conclusion not to adopt it. Recall came in below their own tool, and at a matched context budget it tied plain recency ordering. Cost was much lower, so the honest reading is cheaper but no better. The repo calls that its most credible row, which is a telling thing for a Jev catalog to say. I like that framing because a rejection from a team that already had a working alternative carries more signal than an integration nobody measured. A second row shows no-mistakes removing its Jev review pre-brief after measuring twice: more billed input, essentially no wall-clock gain. The limitation is worth stating. The catalog verifies that cited lines still exist in a file, not that the code runs or that any performance claim holds. It is a starting point, not a verdict. For me the useful idea is simple: read the row's citations before its star count. The most-starred rows are big host projects where Jev is one integration among many. How does your team decide when a promising tool is not worth adopting? #DataEngineering #MachineLearning #OpenSource #AIAgents Read more: https://lnkd.in/e--FNeyj
To view or add a comment, sign in
-
Just got off the waitlist for TypeSafe AI’s early access to #Jev, and after running a few simple tests integrating it into my local Claude Code workflow, the results are on another level. Building robust multi-agent orchestration loops usually means fighting with parsers to get reliable JSON. Jev completely flips that paradigm. It is a pure System 1 model: no text generation, no chatting. You define the exact shape of the answer, and it delivers 100% type-correct structured outputs every single time. A few standout features from my initial runs: 🚀 Speed & Scale: The parallel sampler evaluates multiple questions simultaneously with sub-second latency. It's easily 20–200× faster for routing and classification tasks. 🎯 Calibrated Confidence: Every decision returns a highly accurate probability score. If the model is uncertain, it’s trivial to write conditional edges that fall back to a heavier reasoning model. 💸 Efficiency: It handles instinctive, programmatic judgments at a fraction of the cost (40–1,000× cheaper). It won't replace large reasoning models for complex system design, and it doesn't know niche domains out of the box. But for high-speed tool-calling, stateful execution graphs, and fast logic routing, this architecture is a massive leap forward. Excited to push this further and see how it handles heavier workloads! TypeSafe AI cookbook link: https://lnkd.in/gnh3uCkk
To view or add a comment, sign in
-
-
Qwen 3.8 27b on TERNARY Transformer Architecture is changing game. QAT for TQ_2 quant to read here . Advanced technology ML Read here. https://lnkd.in/g6j5qAWe
Z.ai's latest open-weight model, GLM 5.3 Flash, is now available on Databricks! GLM 5.3 Flash joins 30+ open-source and frontier models on Databricks. On Databricks’ OfficeQA Pro v2 benchmark, it delivers 10% higher quality than GLM-5.2 at just one-tenth the cost, pushing out the quality-cost-pareto frontier. With new multimodal support, it can now read figures and verify web pages at coding. Run GLM 5.3 Flash where your data already lives — governed, secure, and ready for custom AI apps and agents. Control access, spend, and observability across all your AI with Unity Gateway. https://lnkd.in/gyG8433E
To view or add a comment, sign in
-
The Lakeflow Designer August release adds the Unique operator, expands the Aggregate, Select, and Python operators, and improves the canvas, preview, and version-history experience. Anax Holdings (Private) Limited is Authorized Partner for Databricks in #Srilanka #maldives - Reach us out for a comprehensive Data Discovery 💡 Workshop we’re our experts Professionals do Business Discovery workshop FREE.
Z.ai's latest open-weight model, GLM 5.3 Flash, is now available on Databricks! GLM 5.3 Flash joins 30+ open-source and frontier models on Databricks. On Databricks’ OfficeQA Pro v2 benchmark, it delivers 10% higher quality than GLM-5.2 at just one-tenth the cost, pushing out the quality-cost-pareto frontier. With new multimodal support, it can now read figures and verify web pages at coding. Run GLM 5.3 Flash where your data already lives — governed, secure, and ready for custom AI apps and agents. Control access, spend, and observability across all your AI with Unity Gateway. https://lnkd.in/gyG8433E
To view or add a comment, sign in
-
The cheaper model is supposed to be the worse model. Z.ai just shipped one that isn't. GLM-5.3-Flash went out on August 26, 2026 with open weights. It costs one tenth of what GLM-5.2 costs and it scores higher on the benchmarks Z.ai published. 1. DeepSWE v1.1 went from 46.2 to 63.4. That is a 17.2 point jump between two versions from the same lab, run under the same mini-swe-agent harness at 400K context. 2. Terminal-Bench 2.1 went 81.0 to 84.3. Not a leap. But it moved up while the price moved down by 10x, and those two things almost never happen in the same release. 3. It is 320 billion parameters that activates 18 billion per token. You pay for 18B of compute per token and get behaviour from a model an order of magnitude larger. 4. One million tokens of context. Text, image and video input, on a base pretrained over 30 trillion multimodal tokens. 5. The weights are MIT. Not a bespoke licence with a revenue clause attached - MIT. Download it, fork it, ship it commercially. 6. Z.ai's own scorecard: 48.8 on AutomationBench, 78.4 on Toolathlon Verified, 55.3 on Humanity's Last Exam with tools. The story here is not that an open model got better. It is that the price you budgeted against three weeks ago is now the wrong number, and nothing you did caused it to change. If a 10x price cut lands in the middle of your quarter, does your cost model absorb it or does it just quietly become fiction? #OpenSource #AI #LLM #MachineLearning #TechStrategy #Automation #Benchmarks
To view or add a comment, sign in
-
-
⚡ Pareto (SN10) miners make Qwen3-8-27B 3.5x faster in seven days ⚡ Pareto miners reportedly improved inference speed for Qwen3-8-27B by approximately 3.5x in one week, showing how quickly competitive optimization can improve AI model performance. 🔑 Key points 🔹 3.5x speed improvement: Miners significantly increased the model’s throughput over a seven-day optimization period. 🔹 Same model, better execution: The gains came from improving how the model runs rather than simply replacing it with a larger system. 🔹 Hardware and software both matter: Optimization may involve kernels, quantization, batching, memory management, compilation, and GPU utilization. 🔹 Faster inference lowers costs: More tokens per second can reduce the compute required for each request. 🔹 Latency improves user experience: Faster responses are especially valuable for coding agents, chatbots, real-time applications, and automated workflows. 🔹 Competitive mining creates pressure: Miners must continuously improve performance to remain competitive for rewards. 🔹 Benchmark design matters: Speed gains are meaningful only if output quality, accuracy, context length, and reliability remain consistent. 🔹 Optimization may be hardware-specific: A technique that performs well on one GPU may not provide the same improvement on another. 🔹 Sustaining the gain is the next challenge: A short benchmark improvement must translate into reliable production inference. 🔎 Why it matters 🔹 AI infrastructure is often constrained by inference cost and latency rather than model capability alone. 🔹 A decentralized network can use competition to discover optimization techniques that a centralized team may overlook. 🔹 Faster open-model inference could make advanced AI more affordable and expand access to smaller developers. 🔹 The real test is whether the 3.5x improvement survives independent validation, different workloads, and production conditions. 🎯 Bottom line: Pareto’s reported 3.5x improvement shows that model efficiency can advance rapidly when miners compete on execution quality. The achievement is meaningful, but the gain must be measured against accuracy, hardware cost, stability, and real customer workloads before it can be considered a lasting infrastructure breakthrough. #Bittensor #Pareto #SN10 #Qwen3 #AIInference #ArtificialIntelligence #MachineLearning #GPUCompute #DecentralizedAI #AIOptimization #OpenSourceAI #TAO https://lnkd.in/e3ZHEYEN
To view or add a comment, sign in
-
Had some fun this week building with #Jev the new model from TypeSafe AI. Unlike existing LLMs, Jev returns typed judgments & probabilities instead of free text. This means you can't chat with it like you would an LLM chatbot, but you can use it inside software & automations to evaluate well defined inputs and outputs. I built a small web app which models triaging faults in a car. The app reads the customer's description of the problem ("it sort of groans before it catches"). Then it asks the follow-up questions a mechanic would, and ends at a diagnosis with a fix checklist. Results were pretty impressive, when ran independently on simple cases it was able to find the right diagnosis 15/15 times. When run on more complex cases (multiple faults in one session) it found all faults in 14 out of 20 cases, and got the other 6 partially correct; in total finding 37 out of 39 total faults. Full benchmark results are in the repo. A whole session cost ~$0.0008 in model calls, and around 2 seconds waiting for response. The same tokens on Opus 5 would cost over ~100x more. Next up I'll be trying to build a similar application, but one that looks at real logs to see if it can correctly diagnose faults in more complex, real world cases. Link to the repo is in the comments, bring your own api key : )
To view or add a comment, sign in
-
Here's Andrew's full write-up, including the benchmark table: https://verygood.ventures/blog/genui-division-of-labor-problem/