How Jev Uses Meko to Make Code Review Feedback Stick
TypeSafe AI released Jev (a model that does not write text) on September 15. You send Jev a state object and a set of typed questions, and it sends back decisions: a pick from a list, a rating against levels you described, or the probability that a statement is true.
Every answer carries a probability, so your code can tell a confident answer from a coin flip. LangChain published routing and tool-gating middleware built on it two days after launch.
TypeSafe’s models page says Jev is not fine-tuned with customer data. You shape its behavior through the state, instructions, and criteria fields of each request, and nothing else.
So when a Jev-powered check gets something wrong and a person corrects it, where does that correction go, and how can you ensure the correction is implemented in future results?
In this blog, we walk through:
- Why a decision model correction should live outside the model
- What a system needs to ensure that corrections carry into the next request
- A code review check that stops the system from raising a finding after a reviewer has overruled it, tested against the live service
Understanding the Correction Problem
A correction that cannot be trained into Jev has to be stored somewhere else, found again when a similar case occurs, and included in the next request.
TypeSafe’s docs add two constraints.
- Accuracy falls as the state grows with content unrelated to the decision
- The state plus the longest question has to fit in 32k tokens
You cannot add every rule and every past ruling into every request.
A Jev deployment needs:
- A place to keep the rubric (the guidelines or policy it judges against) and the corrections people make to it
- A way to select the few facts that are relevant to each request
- A record of each decision and its probabilities, so you can see why the system did what it did

Meko provides all three.
Powered by YugabyteDB, Meko is an agent-native context engine for multi-agent AI systems. Each project lives in a datapack with private memories, shared knowledge that every member can search (uploaded documents plus promoted memories), and conversations that store each exchange as a trace.
A datapack owner or maintainer promotes a memory into shared knowledge from the learnings tab or with the memory_promote tool, and promotion is one-way. In the loop from Figure 1, approving a ruling is promotion, and the decision record is a conversation.
The review check below calls Meko’s MCP server at https://mcp.mekodata.ai/mcpmeko-jev-code-review-agent repo with the official MCP Python SDK and an API key from Settings > API Keys. There is no LLM in this loop, so the code decides which tool to call and when. The full program is in the .
Building a Code Review Check That Learns
TypeSafe’s use case map calls this job semantic code linting.
Picture the third pull request this week where the check flags a test fixture for returning a bare {"error": "failed"} instead of the team’s error envelope. On the first one, the reviewer wrote,”fixtures under tests/ are exempt because they never ship.” The check never saw that comment, so the same finding comes back on every pull request that touches a fixture. The reviewer has to dismiss it manually every time.
The unit of review is a hunk: one block of changed lines in a diff, the part between one @@ marker and the next, with the file’s path kept on top. A reviewer’s comment attaches to a hunk, so Jev judges hunks. This is the first hunk in the test diff, exactly as Jev sees it:
--- a/tests/fixtures/payments_stub.py
+++ b/tests/fixtures/payments_stub.py
@@ -12,6 +12,9 @@ def charge(request):
if request.headers.get("x-simulate-failure"):
- return JSONResponse({"ok": False}, status_code=500)
+ return JSONResponse({"error": "failed"}, status_code=500)
return JSONResponse({"ok": True, "charge_id": "ch_test_123"})The path is what makes a ruling about tests/ land on this hunk and not on the same change in src/api/payments.py.
Here is the loop with Jev as the judge and the datapack as its record of past rulings:
- App code loads the guidelines once per run and splits the diff into hunks.
- One Jev request per hunk asks a yes/no question per guideline. Rulings are not part of this request.
- Code routes on probability. At 0.7 and above, the hunk gets a finding. Between 0.3 and 0.7, it is escalated for a closer look by a person or a reasoning model. Below 0.3, the finding is dropped.
- For each finding that is not dropped, code searches the datapack for rulings recorded for that guideline and asks Jev, one ruling at a time, whether the ruling covers the hunk. At 0.7 and above, the finding is exempt. Between 0.3 and 0.7, it goes to a person. Below 0.3, the finding stands.
- The decisions, including any ruling that resolved one, are posted to the conversation with
conversation_add_message. - When a reviewer overrules a finding, the ruling is saved with memory_add, and an owner or maintainer promotes it.

Jev reads state literally, so the code writes each ruling in a fixed shape:
- the rule’s name
- the scope
- the reason
The reviewer’s comment above becomes “Error envelope rule: fixtures under tests/ are exempt, because they never ship.”
The name lets the code find the rulings for one rule. The reason keeps the ruling narrow: in our tests, a ruling without one exempted a debug script that replays real customer traffic, and adding the reason dropped that match from 0.9 to 0.23. A ruling only ever exempts. A stricter rule belongs in the guidelines document.
The part that makes this work is keeping two questions apart. The first request judges the hunk solely against the guidelines. Rulings come in afterward, one question per ruling, and only for findings that were not dropped:
# 1. Judge the hunk against the guidelines. No rulings in this request.
answers = (await jev.system_one({"hunk": hunk, "rules": rules}, qs)).answers
# 2. For a contested finding, ask whether each ruling covers the hunk, one question per ruling.
covers = await asyncio.gather(*[jev.system_one({"hunk": hunk, "ruling": r}, {"covers": COVERS})
for r in rulings])The covers question asks Jev to read the ruling’s scope literally:
COVERS = Noul(
instructions="Does the ruling in `ruling` cover `hunk` exactly as the ruling states its scope?",
criteria=NoulCriteria(true="The hunk is squarely inside the scope the ruling names.",
false="The hunk is outside that scope, only resembles it, or goes further than the ruling allows."))Keeping rulings out of the judging request means a ruling can only change the decision it names.
The guidelines and the rulings also come from separate searches, because knowledgebase_search trims results relative to the best match, and a short ruling that matches a hunk closely can push the guidelines out of the results.
TypeSafe calls the middle band confidence-gated routing, and the 0.3 and 0.7 thresholds are the first thing to tune against findings your team has already judged.
Measuring What One Ruling Changes
A ruling should change the decision it was written for and nothing else. We ran this loop against a live Meko datapack and jev-1.13.0 with three guidelines and four hunks, reviewed each hunk three times, promoted the fixtures ruling, and reviewed each hunk three more times.
| Hunk | Before the ruling | After the ruling |
tests/fixtures/payments_stub.py returns a bare {"error": "failed"} | Error envelope finding (0.96 to 0.97) | Exempt (covers 0.87 to 0.89) |
src/api/payments.py returns the same bare error | Error envelope finding (0.97 to 0.98) | Finding stands (0.97 to 0.98) |
tests/fixtures/payments_stub.py logs the request body at INFO | Logging finding (0.98) | Finding stands (0.98) |
tests/helpers/http.py, imported by a production replay script, returns a bare error | Error envelope finding (0.94 to 0.95) | Sent to a person (covers 0.49 to 0.55) |
The ruling resolved the finding it was written for, left the same mistake in production code flagged, and did not touch a different rule in the same file.
The shared helper is the case a person should decide, and it landed in the middle band.
We also checked the cutoffs against 21 cases labeled before the run:
- All 16 clear cases landed in the band we labeled
- Of the five borderline cases:
- Two landed in the review band
- Two stayed findings
- One (the replay script) was exempted until the reason was added to the ruling
Across runs, the probabilities moved by 0.03 at most, in line with TypeSafe’s self-consistency cookbook.

The Same Loop in a Guardrail
Jim Bennett at Arize built a real-time guardrail on Jev for a car dealership chatbot, where each check ends in allow, review, or block. In his demo code, the desk manager’s answer to a review is not stored, so the next similar reply goes to review again.

We ran his reply checks unchanged on jev-1.13.0 and added one desk manager ruling: “Holding a vehicle for a customer for up to a week, with no price stated, is not a price commitment. Allow it.” A reply that held a Tahoe until Saturday went from review to allow (covers 0.86). The same hold at $59,000 stayed blocked, and so did the $1 attack.
A ruling never overrides a block, and the lookup only runs on the review path.
Other Places the Same Loop Fits
The loop works anywhere a team has a written standard, and people overrule the checks.
The rulebook goes into Shared Knowledge. Jev answers narrow questions against it, and a person’s ruling becomes a memory that the next decision can use. TypeSafe’s use case map lists most of these jobs:
- Customer support: Jev classifies issue type and urgency and checks a drafted reply against the refund and escalation policy. A support lead’s ruling, such as “this phrasing is a cancellation threat, not a billing question,” carries into the next ticket.
- Insurance claims: Jev flags missing documents and fraud indicators against the coverage rules. An adjuster’s decision on a claim that Jev sent to review becomes the ruling for similar claims.
- Financial crime: Jev asks whether two records are the same entity and whether a transaction matches a known typology. An investigator’s finding that two names are one customer stops the same alert from firing again.
- Legal and compliance: Jev checks whether a required clause is present or a prohibited claim appears. Counsel’s reading of an ambiguous clause is recorded once and applied whenever the clause recurs.
- Content moderation: Jev scores a post against the community’s own standard. A moderator’s reversal narrows how that rule applies from then on.
- Agent guardrails: Jev allows, reviews, or blocks an agent’s replies and tool calls, and a manager’s ruling resolves the review band, as in the dealership example above.
- Agent memory hygiene: Before an agent writes to Meko, Jev asks whether the text contains a credential, and before memory_promote, whether the promotion was authorized. The rulings here are your team’s own exceptions.
In regulated cases, the trace matters as much as the decision. An auditor can see the questions, probabilities, the policy text that was in state, and the ruling that applied.
Limits to Plan Around
Jev is new, and TypeSafe’s models page says rate limits can change without notice.
Jev reads negations at face value, does not count reliably, and reads dates as text, so does arithmetic and date comparisons in code.
Sending diffs or memories to Jev sends them to a third party. TypeSafe’s legal page lists a data processing agreement and a commitment not to train on user data.
Try It Yourself
The whole loop takes one datapack and one pull request that your team has already argued about. Run the review, promote the ruling from that argument, and run the review again.
- Request Jev early access and generate a key in the TypeSafe console
- Sign up free at mekodata.ai, create a datapack, upload your guidelines with one rule per paragraph, and generate an API key under Settings > API Keys
- Clone the meko-jev-code-review-agent repo and follow its README, which runs the whole loop in five commands
- Read the docs and join the Yugabyte Discord server