← IV Desk / API
Tokens

Drive IV Desk from your own code

Statistics in, never rows. Your loan-level file is parsed and screened in the browser (or in your own process); the API receives only the screen's per-feature statistics in facts: IV, PSI, missing shares, null-IV margins, correlated pairs, row and bad counts, and the flags. No borrower row, no value of any column, ever leaves. The model writes a review: it changes nothing, drops nothing from your data and builds no model.

Everything the web page's paid step does is available over HTTP. Run the variable screen (abnormal months, missing rate, IV, PSI, a null-IV noise screen and a correlation drop), send its results, and get the same review back: a verdict (ready, revise, blocked), a one-sentence headline, a short TL;DR, one review per screening step, a keep / drop / investigate call per feature, leakage suspects with the check that would clear each one, threshold advice tied to your sample, next steps, a cleaning memo a model validator could file, and one answer per browser flag. The natural use is a pipeline hook: screen each new development sample, send the statistics, and file the memo next to the screen.

The easiest way to get a correct body is the page itself: load and screen your file on iv-desk.skillsafe.ai, then press Run input .json under the screen. It is free (no account, no run) and downloads exactly the body the page would send, so you can POST it as it stands. To build it without a browser, run the page's own engine, ivkit.js, in Node (see building the body).

One lane: the task field

Every request names its lane in task. IV Desk has exactly one lane, "review", so the full lane list is that single value. Always send it.

taskwhat it is forreply
reviewA senior credit-risk modeler's review of a finished variable screen: which step results hold for this sample, which features to keep, drop or investigate, what looks like target leakage, which thresholds should change, and the cleaning memo.The eleven keys of the output contract, starting with "task":"review".

If task is missing or different, the model still does the review and writes "review" in the reply's task; the page's checks mark such a reply as a lane mismatch. Do not rely on that. The worked example below is a complete review request and its real reply.

Input fields

The body is one flat JSON object, built in the page by IvKit.buildInput. Every value is a string. Required: task and facts.

fieldtyperequiredwhat it holds
taskstringyesAlways "review"; see the lane.
datasetstringnoA name for the sample ("retail-instalment-2025"). Trimmed and cut at 120 characters. The page sends "" when empty.
contextstringnoFree text: product, target definition, observation and performance windows, which organization is held out. At most 4,000 characters: longer text is cut on a word and ends with [...cut: N more characters not sent], and the model is told the marker may appear. The page sends "" when empty.
factsstringyesThe screen's statistics as a JSON-ENCODED STRING (the output of JSON.stringify), never a nested object. The keys are listed below. Detailed features are capped at 150; the rest are summarized by the step counts.
questionstringnoYour own question, answered inside tldr as one bullet starting "Answer:". At most 2,000 characters, cut the same way as context. The page leaves the key out when there is no question.
retry_notestringnoOnly on a reformat retry, after a reply that could not be parsed: say what was wrong. Never on a first run; see truncation and retries.

The facts string

In the web page, facts is the free screen's result, serialized by IvKit.facts before you pay for anything. Parse it and you get one object with these keys:

keywhat it holds
roles{date, label, org, keys}: the application-date column, the 0/1 label column, the organization column ("" for one organization) and the key columns used to drop duplicates.
settingsEvery threshold the screen used; defaults in the next table. threshold_advice in the reply may only name these keys and must quote their current values.
samplerows_in, invalid_label, invalid_date, duplicates, rows_valid, oos_rows (out-of-sample organizations), modeling_rows_before_months, modeling_rows, modeling_bads, modeling_bad_rate, months_total, months_dropped.
organizationsUp to 40 entries {org, type, rows, bads, bad_rate, months}; type is modeling or oos.
monthsUp to 60 entries {ym, rows, bads, dropped, reason} over the modeling sample; ym is YYYYMM and reason says which abnormal-month rule dropped it ("bads 7 < 10"), else "".
step_countsFeatures flagged by each step: constant, missing, iv, psi, noise, corr. A feature can be flagged by several steps.
features_total, features_detailedFeatures screened, and how many are detailed in features (at most 150).
featuresOne object per detailed feature, flagged ones first, then survivors, then the strongest dropped: name; type (numeric or categorical); missing (share 0-1); iv (overall, 4 decimals, absent for a constant column); drops ("kept", or the flagging steps joined by +, e.g. "iv+noise"). When present: low_iv_orgs (count of organizations under org_iv), psi_max, psi_unstable_orgs (count), null_margin (overall IV minus the largest IV under permuted labels), top_corr ("<feature> <r>"), sentinel_hits (values replaced as missing).
survivors, survivor_countThe features no step flags (up to 200 names) and their count.
correlated_pairsUp to 40 pairs {a, b, r} with absolute Pearson r above max_corr.
noise_screen_ranfalse when the modeling sample has fewer than 1,000 rows (the reference skips the step below that).
flagsEvery flag: id (F1..), severity (high, medium, low), category (sample, months, missing, iv, psi, noise, corr, leakage), message and up to 12 features.
browser_verdictThe screen's hint: blocked (a high sample flag, or nothing could be screened), revise (any other high flag or any medium flag), else ready. The model's verdict is never looser unless it dismissed the flags that set it.
method_notesWhere the browser port differs from the Python reference pipeline: tree bins for IV, the null-IV noise screen standing in for LightGBM null importance, the IV tie-break in the correlation drop, and the PSI binning.

The settings keys and the defaults (IvKit.DEFAULTS):

settingdefaultrule
sentinels"-1, -999, -1111"Numeric values treated as missing.
min_ym_bad, min_ym_total10, 500A modeling month with fewer bads or fewer rows is dropped as abnormal.
missing_ratio0.6Overall missing share above this flags the feature.
overall_iv, org_iv, max_low_iv_orgs0.1, 0.1, 2Overall IV below overall_iv, or IV below org_iv in at least max_low_iv_orgs organizations, flags the feature.
psi, psi_months_ratio, psi_max_orgs, psi_min_month0.1, 0.3333, 6, 100Month-over-month PSI inside each organization; an organization is unstable when PSI exceeds psi in at least that share of its months (months with fewer than psi_min_month non-missing values are skipped); the feature is flagged when at least psi_max_orgs organizations are unstable.
null_perms10Label permutations for the noise screen (at most 50); a feature whose IV does not beat every permuted IV is noise.
max_corr, top_n_keep0.9, 20Of each pair above max_corr the lower-IV feature is flagged, unless both sides are among the top top_n_keep features by IV.

The flags drive the reply. Every flag must come back exactly once in prescan_responses, every survivor (up to 30) gets a call, and a confirmed leakage flag must be carried into leakage_suspects or a non-keep call. If you screen with your own pipeline (toad, scorecardpy, SAS), send the same keys: flags may be [], in which case keep browser_verdict at "ready" so it sets no floor your statistics do not justify.

Building the body

The surest way to match the page is the Run input .json button. The next surest is to run the page's own engine: ivkit.js runs unchanged in Node via require(). parseCsv reads the file, analyze runs the screen (roles are detected from column names when you pass none), and buildInput clips the text fields and serializes the facts:

analyze optionwhat it holds
roles{date, label, org, keys, exclude}; omit to auto-detect (IvKit.detectRoles). exclude lists columns not to screen.
oosOut-of-sample organizations, comma-separated ("partner_d"); they are counted but not screened.
settingsAny of the thresholds above; missing keys take the defaults.
// make-body.js - build the run body with the SAME engine the web page uses.
// Save https://iv-desk.skillsafe.ai/ivkit.js next to this file.
// Usage: node make-body.js loans.csv
const fs = require("fs");
const K = require("./ivkit.js");

const table = K.parseCsv(fs.readFileSync(process.argv[2] || "loans.csv", "utf8"));
const A = K.analyze(table, {
  oos: "partner_d",              // held-out organizations, or ""
  settings: {}                   // thresholds; {} = IvKit.DEFAULTS
});
const body = K.buildInput(A, {
  dataset: "retail-instalment-2025",
  context: "Unsecured instalment loans. Target = 1 if 60+ days past due " +
           "within 12 months of booking.",
  question: ""                   // non-empty: answered in tldr as "Answer: ..."
});
fs.writeFileSync("body.json", JSON.stringify(K.mustBeObject(body)));
console.log(A.flags.length, "flags; browser verdict", A.browser_verdict);
console.log("Idempotency-Key: iv-desk:review:" + K.hashInput(body) + ":a1");

Run on the page's Retail instalment loans example (with that example's own context and question), this reproduces the worked request below byte for byte: hash 8bcf46894487e343. mustBeObject only checks that the body is a plain object; validate the fields yourself as in step 4. From another language, send the same field names and build facts with the keys above.

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}

The token is minted for this app (the guest endpoint takes {"slug":"iv-desk"} in its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….

The input object IS the request body. There is no {"input": …} wrapper. A wrapped body is answered with an unknown field 'input' warning, and the model never sees your screen.

Error codes

statuscodewhat to do
400validation_errorA field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object.
401unauthorizedThe token is missing, malformed or expired. Get a new one from the token page.
402payment_requiredThe balance is below min_credits. Call /estimate first and top up.
403forbiddenThe token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token.
404not_foundUnknown job id, or the app slug does not exist.
409conflictThe same Idempotency-Key was replayed with a different body. Change the key or send the original input.
429rate_limitedToo many requests. Back off and retry; do not tight-loop.
5xxinternalA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. A tiny client

One helper that sends the token, unwraps data and raises on ok: false. The token comes from the token page (Copy token or Copy shell export); step 2 covers the kinds of token and minting one from code.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="iv-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://iv-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

2. Get a token

The easiest route is the token page: it shows the token this browser already holds for iv-desk, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. A guest token, minted with POST /guest and {"slug":"iv-desk"}, can call /me and /estimate; the review is metered, so /run and /run-stream need a personal token.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://iv-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"iv-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}

3. Check the session and the balance

call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}

4. Price the run (free)

/estimate returns the model binding and the credits a run would reserve. It creates no job and charges nothing. Expect model_alias gpt-terra. hold_credits is a reservation, not the price: it is held against your balance while the run executes and released afterwards. min_credits is the least balance that can start a run. What you actually pay is charged_credits, reported on the finished job and in the done event, and it is usually far lower than the hold. The price depends on the size of facts, so a wide file costs more than a narrow one.

The body is the input object itself, with no {"input": …} wrapper. /estimate does not validate the body, so check the shape yourself: an object whose every value is a string, task equal to "review", and facts non-empty and parsing to an object.

# body.json is the input object itself - no {"input": ...} wrapper. Download it with
# the page's free "Run input .json" button, or build it with make-body.js above.
# estimate does not validate it, so check the shape first:
python3 - <<'EOF'
import json
b = json.load(open("body.json"))
assert isinstance(b, dict) and "input" not in b, "no {input: ...} wrapper"
assert all(isinstance(v, str) for v in b.values()), "every value is a string"
assert b.get("task") == "review", "the only lane is review"
assert b.get("facts", "").strip(), "facts is required"
assert isinstance(json.loads(b["facts"]), dict), "facts is a JSON string of an object"
EOF
INPUT=$(cat body.json)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
#   "markup_bps":...,"hold_credits":...,"min_credits":...,
#   "sponsor_enabled":false,"warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.

5. Run it, then poll

POST /run returns a job_id; poll GET /jobs/{id} until it is terminal. The reply is a string at data.output.output: JSON.parse it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the attempt number, iv-desk:review:<hash>:a<attempt> (for example iv-desk:review:8bcf46894487e343:a1), so a retried request returns the same job instead of billing a second run. Use one key per distinct input: a rescreened file, changed settings or a new question is a new hash, and replaying an old key with a different body is a 409. The page uses IvKit.hashInput(body) for the hash (make-body.js prints that key); any stable digest of the body works from other languages. Leave retry_note out of the hash and bump the attempt instead.

# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
# One lane, so the key is iv-desk:review:<hash>:a<attempt>.
KEY="iv-desk:review:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"task\":\"review\",\"verdict\":\"revise\", ...}"},
#   "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c \
  'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json

6. Or stream it

POST /run-stream takes the same body and headers and answers with server-sent events: job (the job id), delta (chunks of the reply) and done (the status, charged_credits, truncated and, when present, the full output). A browser page may receive only tick heartbeats and then done, never a delta, so take the reply from done.output.output when it is there, fall back to the concatenated deltas, and fall back again to GET /jobs/{id}.

# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"task\":\"review\",\"verdict\":\"revise\",\"headline\":\"This"}
# event: done   {"status":"succeeded","charged_credits":...,"truncated":false}

7. Parse the reply

The reply is one JSON object, delivered as a string in data.output.output; you must JSON.parse it. The model is told to send no code fences, but tolerate them: strip a leading ```json and a trailing ```, keep everything from the first { to the last }, and parse that outer object. The keys are the same on every run; see the output contract.

# reply.json holds data.output.output from step 5. Strip any fence, keep the object:
python3 - <<'EOF'
import json, re
t = open("reply.json").read().strip()
t = re.sub(r"^```(?:json)?\s*", "", t, flags=re.I)
t = re.sub(r"\s*```\s*$", "", t)
r = json.loads(t[t.index("{"):t.rindex("}") + 1])
print(r["task"], r["verdict"], "-", r["headline"])
for s in r["step_reviews"]:
    print("STEP", s["step"], s["verdict"], "|", s["note"])
for c in r["feature_calls"]:
    print("CALL", c["feature"], c["call"], "(screen:", c["engine"], "iv", c["iv"], ")")
for s in r["leakage_suspects"]:
    print("LEAK?", s["feature"], "| check:", s["check"])
for a in r["threshold_advice"]:
    print("SET", a["parameter"], a["current"], "->", a["suggested"])
EOF

Invariants worth asserting

The web page holds every reply to the screen before it shows it (recon.js, reconcile()) and lists any disagreement next to the review. Do the same before you file a memo:

# assert-reply.py - the page's reconciliation checks (recon.js), for body.json
# and reply.json from the steps above.
import json
body = json.load(open("body.json")); facts = json.loads(body["facts"])
t = open("reply.json").read(); r = json.loads(t[t.index("{"):t.rindex("}") + 1])
KEYS = ["task", "verdict", "headline", "tldr", "step_reviews", "feature_calls",
        "leakage_suspects", "threshold_advice", "next_steps", "memo", "prescan_responses"]
STEPS = ["sample", "months", "missing", "iv", "psi", "noise", "corr"]
missing = [k for k in KEYS if k not in r]
assert not missing, "keys missing: " + ", ".join(missing)
assert r["task"] == "review", "task is " + str(r["task"])
assert r["verdict"] in ("ready", "revise", "blocked")
# Steps: exactly one review per step, in order.
assert [s["step"] for s in r["step_reviews"]] == STEPS, "one review per step, in order"
assert all(s["verdict"] in ("agree", "adjust", "question") for s in r["step_reviews"])
# Flags: each answered exactly once, none invented.
flags = {f["id"]: f for f in facts["flags"]}
answered = [p["ref"].upper() for p in r["prescan_responses"]]
assert sorted(answered) == sorted(flags), "every flag answered exactly once"
assert all(p["verdict"] in ("confirmed", "dismissed") for p in r["prescan_responses"])
resp = {p["ref"].upper(): p for p in r["prescan_responses"]}
# Features: real names, the screen's keep/drop and IV quoted exactly.
feats = {f["name"]: f for f in facts["features"]}
known = set(feats) | set(facts["survivors"]) | {n for f in facts["flags"] for n in f["features"]}
calls = {c["feature"]: c for c in r["feature_calls"]}
named = list(calls) + [s["feature"] for s in r["leakage_suspects"]]
assert all(n in known for n in named), "a feature not in your file"
for c in r["feature_calls"]:
    assert c["call"] in ("keep", "drop", "investigate")
    f = feats.get(c["feature"])
    if not f:
        continue                      # named in a flag but not detailed
    assert c["engine"] == ("kept" if f["drops"] == "kept" else "dropped"), c["feature"]
    if c["iv"] is not None:
        assert "iv" in f and abs(c["iv"] - f["iv"]) <= 0.00015, "IV misquoted: " + c["feature"]
# Survivors covered; a leakage suspect is never kept.
assert all(n in calls for n in facts["survivors"][:30]), "a survivor has no call"
suspects = {s["feature"] for s in r["leakage_suspects"]}
assert not [n for n in suspects if calls.get(n, {}).get("call") == "keep"], "suspect kept"
# Confirmed high/medium leakage flags carried (or explicitly cleared in the note).
for fid, f in flags.items():
    p = resp[fid]
    if f["category"] != "leakage" or f["severity"] == "low" or p["verdict"] != "confirmed":
        continue
    for n in f["features"]:
        kept = calls.get(n, {}).get("call") == "keep"
        assert n in suspects or (n in calls and not kept) or (kept and n in p["note"]), n
# Threshold advice names a real setting and quotes its current value.
S = facts["settings"]
for a in r["threshold_advice"]:
    assert a["parameter"] in S, "no such setting: " + a["parameter"]
    cur, real = a["current"], S[a["parameter"]]
    if isinstance(real, str):
        assert str(cur).replace(" ", "") == real.replace(" ", ""), a["parameter"]
    else:
        assert abs(float(cur) - real) <= max(1e-6, abs(real) * 0.002), a["parameter"]
# Verdict: never looser than the flags left standing.
rank = {"ready": 0, "revise": 1, "blocked": 2}
floor = 0
for fid, f in flags.items():
    if resp[fid]["verdict"] == "dismissed":
        continue
    lvl = 2 if f["severity"] == "high" and f["category"] == "sample" else 1 if f["severity"] != "low" else 0
    floor = max(floor, lvl)
if not facts["survivors"]:
    floor = 2
assert rank[r["verdict"]] >= floor, "verdict looser than the flags left standing"
overrides = [c["feature"] for c in r["feature_calls"] if c["call"] == "keep" and c["engine"] == "dropped"]
print("ok:", r["verdict"], len(r["feature_calls"]), "calls;", "overrides:", overrides or "none")

The worked example passes every one of these checks, and the page's own reconcile() reports zero disagreements for it.

The output contract

The reply is one JSON object with eleven keys, always all present, in this order. Arrays may be empty ([], never a filler such as "None"). An enum is written "a|b|c": the reply carries exactly one of the values. Text fields are plain prose, no Markdown, no emoji, each under 400 characters except memo paragraphs (up to 900). Every feature name and every number comes from facts.

{"task":"review",
 "verdict":"ready|revise|blocked",
 "headline":"...",
 "tldr":["...","..."],
 "step_reviews":[{"step":"sample","verdict":"agree|adjust|question","note":"..."}],
 "feature_calls":[{"feature":"...","call":"keep|drop|investigate",
                   "engine":"kept|dropped","iv":0.1234,"reason":"...","ref":"F2"}],
 "leakage_suspects":[{"feature":"...","evidence":"...","check":"..."}],
 "threshold_advice":[{"parameter":"psi_max_orgs","current":6,"suggested":2,"why":"..."}],
 "next_steps":["...","..."],
 "memo":["...","..."],
 "prescan_responses":[{"ref":"F1","verdict":"confirmed|dismissed","note":"..."}]}
keyshapewhat it holds
taskstringAlways "review".
verdictenumCan this screen go to binning and modeling; see the next table.
headlinestringOne sentence saying whether the screen can go to modeling.
tldrarray of strings2-5 bullets. When you sent a question, one bullet starts "Answer:".
step_reviewsarray of {step, verdict, note}Exactly seven, in order: sample, months, missing, iv, psi, noise, corr. agree: the rule and its result fit the sample. adjust: a threshold should change (given in threshold_advice). question: the result cannot be trusted as run, for example a rule that could not fire as set.
feature_callsarray of {feature, call, engine, iv, reason, ref}At most 40: every survivor (up to 30), every feature of a confirmed flag, and any dropped feature worth bringing back. engine is what the screen decided; iv is copied from facts or null; ref is the flag id the call answers, or "". A keep on an engine dropped feature is an override and its reason names the step.
leakage_suspectsarray of {feature, evidence, check}Features that may encode the outcome, with the evidence from facts and the concrete check that would clear or convict each one. May include a feature the browser did not flag, saying so in evidence.
threshold_advicearray of {parameter, current, suggested, why}At most 6. parameter is a key of facts.settings, current its value there; suggested is a number, or a comma-separated list for sentinels.
next_stepsarray of strings3-7 concrete actions, in order.
memoarray of strings2-4 paragraphs a model validator could file as the cleaning report: the sample, what each step removed, the leakage and stability concerns, the survivors and what happens next.
prescan_responsesarray of {ref, verdict, note}Exactly one per flag in facts.flags: confirmed (a real issue for this sample) or dismissed (the note names the fact that shows it is harmless). Empty when no flags were sent.

Verdict

verdictmeaning
blockedThe sample cannot support a variable screen yet: a confirmed high flag of category sample (too few bads, nothing screened, nothing survives), or no survivors.
reviseUsable, but something must change before modeling: any other confirmed high flag (leakage) or any confirmed medium flag.
readyNothing confirmed at high or medium; the survivors can go to binning and modeling.

The verdict floor: the model may be stricter than facts.browser_verdict, never looser, unless it dismissed the flags that set it. Your own check computes the floor from the flags left standing, exactly as in the invariants.

Enums

wherevaluesthe page's fallback
verdictready, revise, blockedrevise
step_reviews[].stepsample, months, missing, iv, psi, noise, corr (in this order)kept as written
step_reviews[].verdictagree, adjust, questionquestion
feature_calls[].callkeep, drop, investigateinvestigate
feature_calls[].enginekept, dropped""
prescan_responses[].verdictconfirmed, dismissedconfirmed

Worked example: review

The page's Retail instalment loans example: 22,273 loan rows from three lenders (bank_a, bank_b, broker_c) plus partner_d, held out of sample; 19,374 modeling rows with 2,122 bads over 12 months. Twenty features were screened and 6 survive. The screen raised 6 flags: F2 high leakage (overdue_days_m3 has IV 16.2826, bureau_score 0.5756), F1 and F5 medium PSI (the rule needs 6 unstable organizations and there are 3, so fraud_score_v2, unstable in all three, was not dropped by PSI), F3 medium missing (sentinels in features that hold other negatives), F6 medium correlation (a pair with r 0.9998 left in place by top_n_keep) and F4 low (a constant column), so browser_verdict is revise. Its Idempotency-Key from the page is iv-desk:review:8bcf46894487e343:a1.

The request body. The facts string is 6,892 characters; only its first 150 are shown here, ending in a marked …. Every other field is shown in full, exactly as sent:

{
 "task": "review",
 "dataset": "retail-instalment-2025",
 "context": "Unsecured instalment loans, three lenders plus one partner held out of sample. Target = 1 if 60+ days past due within 12 months of booking. Features are meant to be known at application time.",
 "facts": "{\"roles\":{\"date\":\"apply_date\",\"label\":\"target\",\"org\":\"org_info\",\"keys\":[\"loan_id\"]},\"settings\":{\"sentinels\":\"-1, -999, -1111\",\"min_ym_bad\":10,\"min_ym_…",
 "question": "Which of these would you take into a logistic scorecard?"
}

The same facts, decoded. Real values throughout; the lines starting with … are trims made for this page, not part of the JSON:

{
 "roles": {"date": "apply_date", "label": "target", "org": "org_info", "keys": ["loan_id"]},
 "settings": {
  "sentinels": "-1, -999, -1111", "min_ym_bad": 10, "min_ym_total": 500,
  "missing_ratio": 0.6, "overall_iv": 0.1, "org_iv": 0.1,
  "max_low_iv_orgs": 2, "psi": 0.1, "psi_months_ratio": 0.3333,
  "psi_max_orgs": 6, "psi_min_month": 100, "null_perms": 10,
  "max_corr": 0.9, "top_n_keep": 20
 },
 "sample": {
  "rows_in": 22273, "invalid_label": 0, "invalid_date": 0,
  "duplicates": 0, "rows_valid": 22273, "oos_rows": 2899,
  "modeling_rows_before_months": 19374, "modeling_rows": 19374, "modeling_bads": 2122,
  "modeling_bad_rate": 0.1095, "months_total": 12, "months_dropped": 0
 },
 "organizations": [
  {"org": "bank_a", "type": "modeling", "rows": 7592, "bads": 818, "bad_rate": 0.1077, "months": 12},
  {"org": "bank_b", "type": "modeling", "rows": 6647, "bads": 655, "bad_rate": 0.0985, "months": 12},
  {"org": "broker_c", "type": "modeling", "rows": 5135, "bads": 649, "bad_rate": 0.1264, "months": 12},
  {"org": "partner_d", "type": "oos", "rows": 2899, "bads": 329, "bad_rate": 0.1135, "months": 12}
 ],
 "months": [
  {"ym": "202501", "rows": 1643, "bads": 191, "dropped": false, "reason": ""},
  {"ym": "202502", "rows": 1591, "bads": 176, "dropped": false, "reason": ""}
   … 10 more months, 202503 to 202512, all with dropped false
 ],
 "step_counts": {"constant": 1, "missing": 1, "iv": 13, "psi": 0, "noise": 8, "corr": 0},
 "features_total": 20,
 "features_detailed": 20,
 "features": [
  {"name": "overdue_days_m3", "type": "numeric", "missing": 0, "iv": 16.2826, "drops": "kept", "psi_max": 0.0219, "null_margin": 16.2709},
  {"name": "bureau_score", "type": "numeric", "missing": 0, "iv": 0.5756, "drops": "kept", "psi_max": 0.11, "null_margin": 0.5633},
  {"name": "util_revolving", "type": "numeric", "missing": 0, "iv": 0.14, "drops": "kept", "psi_max": 0.1154, "null_margin": 0.1259, "top_corr": "util_revolving_pct 0.9998"},
  {"name": "util_revolving_pct", "type": "numeric", "missing": 0, "iv": 0.1393, "drops": "kept", "psi_max": 0.117, "null_margin": 0.1226, "top_corr": "util_revolving 0.9998"},
  {"name": "noise_a", "type": "numeric", "missing": 0.0002, "iv": 0.0155, "drops": "iv+noise", "low_iv_orgs": 3, "psi_max": 0.0943, "null_margin": 0, "sentinel_hits": 4},
  {"name": "fraud_score_v2", "type": "numeric", "missing": 0, "iv": 0.0147, "drops": "iv+noise", "low_iv_orgs": 3, "psi_max": 0.3709, "psi_unstable_orgs": 3, "null_margin": -0.0018},
  {"name": "bal_change_3m", "type": "numeric", "missing": 0.0691, "iv": 0.0126, "drops": "iv+noise", "low_iv_orgs": 3, "psi_max": 0.1035, "null_margin": -0.0046, "sentinel_hits": 1339},
  {"name": "const_flag", "type": "categorical", "missing": 0, "drops": "constant"}
   … 12 more features: dti, inq_6m, income_monthly, employment_months, social_score, app_channel, residence_type, age, loan_amount, noise_b, device_os, term_months
 ],
 "survivors": ["bureau_score", "dti", "inq_6m", "util_revolving", "util_revolving_pct", "overdue_days_m3"],
 "survivor_count": 6,
 "correlated_pairs": [
  {"a": "util_revolving", "b": "util_revolving_pct", "r": 0.9998}
 ],
 "noise_screen_ran": true,
 "flags": [
  {"id": "F1", "severity": "medium", "category": "psi", "message": "The PSI rule needs 6 unstable organizations to drop a feature, but the modeling sample has only 3 - the PSI step cannot drop anything as set.", "features": []},
  {"id": "F2", "severity": "high", "category": "leakage", "message": "2 feature(s) have overall IV above 0.5 - bureau_score (0.5756), overdue_days_m3 (16.2826). IV that strong usually means the feature was observed after the outcome or encodes the label.", "features": ["bureau_score", "overdue_days_m3"]},
  {"id": "F3", "severity": "medium", "category": "missing", "message": "Sentinel values were treated as missing in feature(s) that also hold other negative numbers, so a real value may have been blanked: noise_a, bal_change_3m.", "features": ["noise_a", "bal_change_3m"]},
  {"id": "F4", "severity": "low", "category": "sample", "message": "1 constant column(s) dropped before screening: const_flag.", "features": ["const_flag"]},
  {"id": "F5", "severity": "medium", "category": "psi", "message": "Unstable month over month in every modeling organization, yet kept by the PSI rule as set: fraud_score_v2 (max PSI 0.3709).", "features": ["fraud_score_v2"]},
  {"id": "F6", "severity": "medium", "category": "corr", "message": "1 pair(s) above the correlation threshold were left in place because both sides are among the top 20 features by IV: util_revolving ~ util_revolving_pct (r 0.9998).", "features": ["util_revolving", "util_revolving_pct"]}
 ],
 "browser_verdict": "revise",
 "method_notes": [
  "IV bins: best-first decision tree, at most 5 leaves, 1% minimum leaf, missing as its own bin (toad dt recipe, re-implemented).",
  "Noise screen: univariate null-IV test against label permutations - a stand-in for the reference's LightGBM null-importance gain, which does not run in a browser.",
  "Correlation drop keeps the higher-IV feature of each pair (the reference uses LightGBM gain).",
  "PSI: month-over-month within each organization, 10 quantile bins of the earlier month plus a missing bin; values above the earlier month's maximum go to the top bin."
 ]
}

The reply (data.output.output, parsed), as the model returned it. The lines starting with … mark trimmed array entries; every value shown is real:

{
 "task": "review",
 "verdict": "revise",
 "headline": "This screen needs one leakage check before modeling: overdue_days_m3's near-perfect IV must be confirmed as application-time before the survivors go to binning.",
 "tldr": [
  "Answer: dti, inq_6m and bureau_score are safe for the scorecard today; hold overdue_days_m3 until its timing is confirmed, and use only one of util_revolving/util_revolving_pct.",
  "6 survivors from 19,374 modeling rows across 3 organizations (2,122 bads, 10.95% bad rate), with partner_d's 2,899 rows correctly held out of sample.",
  "overdue_days_m3 (IV 16.28, null_margin 16.27) is too strong to trust without a timing check; bureau_score's IV 0.58 is plausible for a bureau score and is cleared.",
  "PSI and correlation both have configuration gaps: psi_max_orgs=6 can never fire with only 3 organizations, and top_n_keep=20 equals the full feature count so the correlation guard never engages.",
  "Sentinel handling likely masks real values in bal_change_3m and noise_a - all of their recorded missingness is sentinel hits, though both are already IV-dropped so it costs nothing here."
 ],
 "step_reviews": [
  {"step": "sample",
   "verdict": "agree",
   "note": "22,273 rows in, no invalid labels/dates, no duplicates. 2,899 partner_d rows held out. Modeling sample is 19,374 rows, 2,122 bads (10.95%) across bank_a/bank_b/broker_c - plenty of bads for a stable screen."},
  {"step": "months",
   "verdict": "agree",
   "note": "12 months, none dropped. Monthly rows run 1,575-1,661 and bads 157-191, comfortably above min_ym_bad=10 and min_ym_total=500 - the calendar is sound."},
  {"step": "missing",
   "verdict": "adjust",
   "note": "The 0.6 cutoff correctly caught only social_score (0.7588 missing). But bal_change_3m's whole missing share (0.0691) and noise_a's (0.0002) equal their sentinel_hits exactly - every 'missing' there is a sentinel replacement, not a true gap; review the sentinel list."},
  {"step": "iv",
   "verdict": "agree",
   "note": "overall_iv=0.1 with org_iv=0.1/max_low_iv_orgs=2 correctly dropped 13 weak features and kept the 6 survivors plus bureau_score and overdue_days_m3, which the leakage check (F2) handles separately."},
  {"step": "psi",
   "verdict": "question",
   "note": "psi_max_orgs=6 exceeds the 3 modeling organizations that exist, so the rule structurally cannot drop a feature (F1); step_counts.psi=0 confirms it never has. fraud_score_v2 is unstable in all 3 orgs (PSI 0.3709) yet survived (F5)."},
  {"step": "noise",
   "verdict": "agree",
   "note": "The null-IV permutation screen correctly caught both deliberately noisy probes, noise_a and noise_b (null_margin ~0 and -0.0126), a good sanity check that the screen works as intended."},
  {"step": "corr",
   "verdict": "question",
   "note": "top_n_keep=20 equals features_total=20, so the 'both in the top N' protection is trivially true for every pair; util_revolving/util_revolving_pct (r=0.9998) were kept together with no real prioritization (F6)."}
 ],
 "feature_calls": [
  {"feature": "overdue_days_m3",
   "call": "investigate",
   "engine": "kept",
   "iv": 16.2826,
   "reason": "IV 16.2826 with null_margin 16.2709 is far beyond what an application-time feature should show; the name suggests a days-overdue window that may overlap the 12-month performance period used to build the target. Hold for a timing check before scorecard use.",
   "ref": "F2"},
  {"feature": "bureau_score",
   "call": "keep",
   "engine": "kept",
   "iv": 0.5756,
   "reason": "IV 0.5756 (null_margin 0.5633) is high but a genuine bureau score can legitimately reach that on its own; no leakage indicator beyond IV strength. Safe for the scorecard.",
   "ref": "F2"},
  {"feature": "util_revolving",
   "call": "keep",
   "engine": "kept",
   "iv": 0.14,
   "reason": "IV 0.14, the stronger side of a near-duplicate pair (r 0.9998 with util_revolving_pct). Keep this one.",
   "ref": "F6"},
  {"feature": "util_revolving_pct",
   "call": "drop",
   "engine": "kept",
   "iv": 0.1393,
   "reason": "Overriding the corr step: it kept this feature only because top_n_keep=20 equals the full feature count, so the correlation guard never actually screens anything here. r=0.9998 with util_revolving means it adds no information; drop the redundant copy.",
   "ref": "F6"}
   … 7 more calls: dti keep, inq_6m keep, noise_a drop, bal_change_3m drop, const_flag drop, fraud_score_v2 drop, income_monthly investigate
 ],
 "leakage_suspects": [
  {"feature": "overdue_days_m3",
   "evidence": "IV 16.2826, essentially unchanged after label permutation (null_margin 16.2709) - far outside the range a legitimate application-time feature should show.",
   "check": "Confirm overdue_days_m3 is computed strictly from data timestamped before apply_date and not from the 12-month performance window; rebuild it from pre-application snapshots only and rescore its IV."}
 ],
 "threshold_advice": [
  {"parameter": "sentinels",
   "current": "-1, -999, -1111",
   "suggested": "-999, -1111",
   "why": "bal_change_3m's missing share (0.0691) equals its sentinel_hits (1339) over modeling_rows exactly, and noise_a's missing (0.0002) equals its 4 sentinel_hits the same way - every recorded 'missing' value in both is a sentinel hit. Confirm -1 is not a legitimate small value before keeping it as a sentinel."},
  {"parameter": "psi_max_orgs",
   "current": 6,
   "suggested": 2,
   "why": "Only 3 modeling organizations exist (bank_a, bank_b, broker_c), so psi_max_orgs=6 can never be reached (F1); fraud_score_v2 is unstable (PSI 0.3709) in all 3 yet survived (F5)."},
  {"parameter": "top_n_keep",
   "current": 20,
   "suggested": 8,
   "why": "top_n_keep=20 equals features_total=20, so the correlated-pair protection always applies to any pair; util_revolving and util_revolving_pct (r=0.9998) were kept together with no real prioritization (F6)."}
 ],
 "next_steps": [
  "Confirm overdue_days_m3 is built from data timestamped strictly before apply_date, not the 12-month performance window; rebuild and rescore if it isn't.",
  "Drop util_revolving_pct and keep util_revolving (r=0.9998, marginally higher IV) before binning.",
  "Lower psi_max_orgs to 2 of 3 organizations and rerun PSI so instability like fraud_score_v2's can actually drop a feature.",
  "Lower top_n_keep well below features_total (e.g. 8) so the correlation guard stops protecting every pair by default.",
  "Review the sentinel list (-1, -999, -1111) against bal_change_3m and noise_a specifically before trusting their missing shares.",
  "Take bureau_score, dti, inq_6m and util_revolving into scorecard binning now; hold overdue_days_m3 pending the leakage check.",
  "Revisit income_monthly after cleaning its sentinel-driven missingness (943 hits) - its IV of 0.0548 is borderline and worth a second look."
 ],
 "memo": [
  "The retail-instalment-2025 sample has 22,273 loan-level rows, no invalid labels or dates and no duplicates. 2,899 rows from partner_d are held out of sample entirely; the modeling sample is bank_a, bank_b and broker_c, 19,374 rows, 2,122 bads (10.95% bad rate) over 12 months, none dropped. Monthly bads run 157-191 and rows 1,575-1,661, comfortably above the min_ym_bad and min_ym_total floors.",
  "Of 20 features, const_flag was removed as constant, social_score failed missing-rate (75.88%), 13 failed IV, 8 failed noise (incl. probes noise_a/noise_b), 0 failed PSI/corr - not for lack of instability or correlation, but because psi_max_orgs=6 exceeds the 3 organizations and top_n_keep=20 equals features_total. fraud_score_v2 and util_revolving_pct both slipped through that way."
   … 2 more paragraphs: leakage and stability concerns, survivors and next steps
 ],
 "prescan_responses": [
  {"ref": "F1",
   "verdict": "confirmed",
   "note": "Only 3 modeling organizations exist against psi_max_orgs=6 - the rule needs more unstable organizations than exist, so PSI can never drop a feature as configured (step_counts.psi=0 confirms it never has)."},
  {"ref": "F2",
   "verdict": "confirmed",
   "note": "overdue_days_m3 (IV 16.2826, null_margin 16.2709) is implausibly strong for an application-time feature and is treated as a leakage suspect. bureau_score (IV 0.5756, null_margin 0.5633) is a standard bureau score that can legitimately reach this IV on its own - cleared, kept as a survivor."},
  {"ref": "F3",
   "verdict": "confirmed",
   "note": "bal_change_3m's missing share (0.0691) exactly equals its sentinel_hits over modeling_rows (1339/19374), so every 'missing' value there is a sentinel replacement, not a true gap; noise_a is the same (4/19374). Both are already IV-dropped so there is no scorecard impact, but the sentinel list should be reviewed."}
   … F4 confirmed, F5 confirmed, F6 confirmed
 ]
}

Note what the review did that the screen could not: it kept bureau_score (a bureau score can reach IV 0.5 on its own) while holding overdue_days_m3 as a leakage suspect, called util_revolving_pct drop although the screen kept it, and flagged the two rules that could not fire as set. It changed nothing: the screen's result is still what engine says.

Truncation and partial results

When the balance sits between min_credits and hold_credits, the run is not refused: it executes with a reduced output cap and reports "truncated": true, in the done event of /run-stream and on the job from GET /jobs/{id}. What you hold is then a prefix of the reply, and it will not parse as it stands. The web page closes the cut-off JSON (Recon.closeJson in recon.js: close an open string, drop a dangling comma or key, close every open array and object, and if that still does not parse, cut back to the previous comma and try again), parses what is left, and shows the sections that arrived as "N of 11 sections recovered". A stream ended early error gets the same treatment on the deltas received so far.

The keys arrive in contract order, so a cut costs the tail first: memo and prescan_responses go first, then next_steps and threshold_advice. A truncated review can therefore carry a verdict while the flag answers that justify it are missing, and it will fail the flag and verdict-floor checks above. Never file one as final: check the flag, top up, and resubmit with the attempt suffix on the Idempotency-Key incremented (iv-desk:review:<hash>:a2).

If a complete reply will not parse as one JSON object, the page retries once, as the next attempt, with a retry_note saying what was wrong and asking for only the JSON object for task review, no prose, no code fences, every key present. Do the same: keep the hash of the original body, bump the attempt, add retry_note.