Drive IV Desk from your own code
Statistics in, never rows. Your loan-level file is parsed and screened in the
browser (or in your own process); the API receives only the screen's per-feature statistics in
facts: IV, PSI, missing shares, null-IV margins, correlated pairs, row and bad counts,
and the flags. No borrower row, no value of any column, ever leaves. The model writes a review: it
changes nothing, drops nothing from your data and builds no model.
Everything the web page's paid step does is available over HTTP. Run the variable screen (abnormal
months, missing rate, IV, PSI, a null-IV noise screen and a correlation drop), send its results, and
get the same review back: a verdict (ready, revise,
blocked), a one-sentence headline, a short TL;DR, one review per screening step, a
keep / drop / investigate call per feature, leakage suspects with the check that would clear each
one, threshold advice tied to your sample, next steps, a cleaning memo a model validator could file,
and one answer per browser flag. The natural use is a pipeline hook: screen each new development
sample, send the statistics, and file the memo next to the screen.
The easiest way to get a correct body is the page itself: load and screen your
file on iv-desk.skillsafe.ai, then press
Run input .json under the screen. It is free (no account, no run) and downloads
exactly the body the page would send, so you can POST it as it stands. To build it
without a browser, run the page's own engine, ivkit.js, in Node (see
building the body).
One lane: the task field
Every request names its lane in task. IV Desk has exactly one lane,
"review", so the full lane list is that single value. Always send it.
| task | what it is for | reply |
|---|---|---|
review | A senior credit-risk modeler's review of a finished variable screen: which step results hold for this sample, which features to keep, drop or investigate, what looks like target leakage, which thresholds should change, and the cleaning memo. | The eleven keys of the output contract, starting with "task":"review". |
If task is missing or different, the model still does the review and writes
"review" in the reply's task; the page's checks mark such a reply as a lane
mismatch. Do not rely on that. The worked example below is a complete
review request and its real reply.
Input fields
The body is one flat JSON object, built in the page by IvKit.buildInput. Every value
is a string. Required: task and facts.
| field | type | required | what it holds |
|---|---|---|---|
task | string | yes | Always "review"; see the lane. |
dataset | string | no | A name for the sample ("retail-instalment-2025"). Trimmed and cut at 120 characters. The page sends "" when empty. |
context | string | no | Free text: product, target definition, observation and performance windows, which organization is held out. At most 4,000 characters: longer text is cut on a word and ends with [...cut: N more characters not sent], and the model is told the marker may appear. The page sends "" when empty. |
facts | string | yes | The screen's statistics as a JSON-ENCODED STRING (the output of JSON.stringify), never a nested object. The keys are listed below. Detailed features are capped at 150; the rest are summarized by the step counts. |
question | string | no | Your own question, answered inside tldr as one bullet starting "Answer:". At most 2,000 characters, cut the same way as context. The page leaves the key out when there is no question. |
retry_note | string | no | Only on a reformat retry, after a reply that could not be parsed: say what was wrong. Never on a first run; see truncation and retries. |
The facts string
In the web page, facts is the free screen's result, serialized by
IvKit.facts before you pay for anything. Parse it and you get one object with these keys:
| key | what it holds |
|---|---|
roles | {date, label, org, keys}: the application-date column, the 0/1 label column, the organization column ("" for one organization) and the key columns used to drop duplicates. |
settings | Every threshold the screen used; defaults in the next table. threshold_advice in the reply may only name these keys and must quote their current values. |
sample | rows_in, invalid_label, invalid_date, duplicates, rows_valid, oos_rows (out-of-sample organizations), modeling_rows_before_months, modeling_rows, modeling_bads, modeling_bad_rate, months_total, months_dropped. |
organizations | Up to 40 entries {org, type, rows, bads, bad_rate, months}; type is modeling or oos. |
months | Up to 60 entries {ym, rows, bads, dropped, reason} over the modeling sample; ym is YYYYMM and reason says which abnormal-month rule dropped it ("bads 7 < 10"), else "". |
step_counts | Features flagged by each step: constant, missing, iv, psi, noise, corr. A feature can be flagged by several steps. |
features_total, features_detailed | Features screened, and how many are detailed in features (at most 150). |
features | One object per detailed feature, flagged ones first, then survivors, then the strongest dropped: name; type (numeric or categorical); missing (share 0-1); iv (overall, 4 decimals, absent for a constant column); drops ("kept", or the flagging steps joined by +, e.g. "iv+noise"). When present: low_iv_orgs (count of organizations under org_iv), psi_max, psi_unstable_orgs (count), null_margin (overall IV minus the largest IV under permuted labels), top_corr ("<feature> <r>"), sentinel_hits (values replaced as missing). |
survivors, survivor_count | The features no step flags (up to 200 names) and their count. |
correlated_pairs | Up to 40 pairs {a, b, r} with absolute Pearson r above max_corr. |
noise_screen_ran | false when the modeling sample has fewer than 1,000 rows (the reference skips the step below that). |
flags | Every flag: id (F1..), severity (high, medium, low), category (sample, months, missing, iv, psi, noise, corr, leakage), message and up to 12 features. |
browser_verdict | The screen's hint: blocked (a high sample flag, or nothing could be screened), revise (any other high flag or any medium flag), else ready. The model's verdict is never looser unless it dismissed the flags that set it. |
method_notes | Where the browser port differs from the Python reference pipeline: tree bins for IV, the null-IV noise screen standing in for LightGBM null importance, the IV tie-break in the correlation drop, and the PSI binning. |
The settings keys and the defaults (IvKit.DEFAULTS):
| setting | default | rule |
|---|---|---|
sentinels | "-1, -999, -1111" | Numeric values treated as missing. |
min_ym_bad, min_ym_total | 10, 500 | A modeling month with fewer bads or fewer rows is dropped as abnormal. |
missing_ratio | 0.6 | Overall missing share above this flags the feature. |
overall_iv, org_iv, max_low_iv_orgs | 0.1, 0.1, 2 | Overall IV below overall_iv, or IV below org_iv in at least max_low_iv_orgs organizations, flags the feature. |
psi, psi_months_ratio, psi_max_orgs, psi_min_month | 0.1, 0.3333, 6, 100 | Month-over-month PSI inside each organization; an organization is unstable when PSI exceeds psi in at least that share of its months (months with fewer than psi_min_month non-missing values are skipped); the feature is flagged when at least psi_max_orgs organizations are unstable. |
null_perms | 10 | Label permutations for the noise screen (at most 50); a feature whose IV does not beat every permuted IV is noise. |
max_corr, top_n_keep | 0.9, 20 | Of each pair above max_corr the lower-IV feature is flagged, unless both sides are among the top top_n_keep features by IV. |
The flags drive the reply. Every flag must come back exactly once in
prescan_responses, every survivor (up to 30) gets a call, and a confirmed leakage flag
must be carried into leakage_suspects or a non-keep call. If you screen
with your own pipeline (toad, scorecardpy, SAS), send the same keys: flags may be
[], in which case keep browser_verdict at "ready" so it sets no
floor your statistics do not justify.
Building the body
The surest way to match the page is the Run input .json button. The next surest is
to run the page's own engine: ivkit.js runs unchanged in Node via
require(). parseCsv reads the file, analyze runs the screen
(roles are detected from column names when you pass none), and buildInput clips the
text fields and serializes the facts:
| analyze option | what it holds |
|---|---|
roles | {date, label, org, keys, exclude}; omit to auto-detect (IvKit.detectRoles). exclude lists columns not to screen. |
oos | Out-of-sample organizations, comma-separated ("partner_d"); they are counted but not screened. |
settings | Any of the thresholds above; missing keys take the defaults. |
// make-body.js - build the run body with the SAME engine the web page uses.
// Save https://iv-desk.skillsafe.ai/ivkit.js next to this file.
// Usage: node make-body.js loans.csv
const fs = require("fs");
const K = require("./ivkit.js");
const table = K.parseCsv(fs.readFileSync(process.argv[2] || "loans.csv", "utf8"));
const A = K.analyze(table, {
oos: "partner_d", // held-out organizations, or ""
settings: {} // thresholds; {} = IvKit.DEFAULTS
});
const body = K.buildInput(A, {
dataset: "retail-instalment-2025",
context: "Unsecured instalment loans. Target = 1 if 60+ days past due " +
"within 12 months of booking.",
question: "" // non-empty: answered in tldr as "Answer: ..."
});
fs.writeFileSync("body.json", JSON.stringify(K.mustBeObject(body)));
console.log(A.flags.length, "flags; browser verdict", A.browser_verdict);
console.log("Idempotency-Key: iv-desk:review:" + K.hashInput(body) + ":a1");
Run on the page's Retail instalment loans example (with that example's own context and
question), this reproduces the worked request below byte for byte: hash
8bcf46894487e343. mustBeObject only checks that the body is a plain
object; validate the fields yourself as in step 4. From another language, send the same field names
and build facts with the keys above.
Base URL and the envelope
Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses
the same envelope, so one helper covers the whole API:
{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}
The token is minted for this app (the guest endpoint takes {"slug":"iv-desk"} in its
body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….
The input object IS the request body. There is no {"input": …}
wrapper. A wrapped body is answered with an unknown field 'input' warning, and the
model never sees your screen.
Error codes
| status | code | what to do |
|---|---|---|
| 400 | validation_error | A field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object. |
| 401 | unauthorized | The token is missing, malformed or expired. Get a new one from the token page. |
| 402 | payment_required | The balance is below min_credits. Call /estimate first and top up. |
| 403 | forbidden | The token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token. |
| 404 | not_found | Unknown job id, or the app slug does not exist. |
| 409 | conflict | The same Idempotency-Key was replayed with a different body. Change the key or send the original input. |
| 429 | rate_limited | Too many requests. Back off and retry; do not tight-loop. |
| 5xx | internal | A server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice. |
1. A tiny client
One helper that sends the token, unwraps data and raises on ok: false.
The token comes from the token page (Copy token or
Copy shell export); step 2 covers the kinds of token and minting one from code.
# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="iv-desk"
TOKEN="$SKILLSAFE_TOKEN" # from https://iv-desk.skillsafe.ai/tokens.html
call() { # call <path> [json-body]
if [ -n "$2" ]; then
curl -sS -X POST "$BASE/$1" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d "$2"
else
curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
fi
}
import json, os, urllib.error, urllib.request
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "iv-desk"
TOKEN = os.environ.get("SKILLSAFE_TOKEN", "YOUR_TOKEN") # from https://iv-desk.skillsafe.ai/tokens.html
def call(path, body=None):
"""Returns the unwrapped `data`, or raises with the API error code."""
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(f"{BASE}/{path}", data=data, method="POST" if body is not None else "GET")
req.add_header("Authorization", f"Bearer {TOKEN}")
if body is not None:
req.add_header("Content-Type", "application/json")
try:
with urllib.request.urlopen(req) as r:
payload = json.load(r)
except urllib.error.HTTPError as e:
payload = json.load(e)
if not payload.get("ok"):
err = payload.get("error", {})
raise RuntimeError(f"{err.get('code')}: {err.get('message')}")
return payload["data"]
import { readFileSync } from "node:fs";
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "iv-desk";
// Paste the token from https://iv-desk.skillsafe.ai/tokens.html into a file named "token",
// or replace the fallback with it.
let TOKEN = "YOUR_TOKEN";
try { TOKEN = readFileSync("token", "utf8").trim(); } catch {}
async function call(path, body) {
const res = await fetch(`${BASE}/${path}`, {
method: body ? "POST" : "GET",
headers: {
Authorization: `Bearer ${TOKEN}`,
...(body ? { "Content-Type": "application/json" } : {}),
},
body: body ? JSON.stringify(body) : undefined,
});
const payload = await res.json();
if (!payload.ok) throw new Error(`${payload.error.code}: ${payload.error.message}`);
return payload.data;
}
package main
import (
"bufio"
"bytes"
"crypto/sha256"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
const (
base = "https://api.skillsafe.ai/v1/app-api"
slug = "iv-desk"
)
var token = os.Getenv("SKILLSAFE_TOKEN") // from https://iv-desk.skillsafe.ai/tokens.html
type envelope struct {
OK bool `json:"ok"`
Data json.RawMessage `json:"data"`
Error struct {
Code string `json:"code"`
Message string `json:"message"`
} `json:"error"`
}
func call(path string, body any) (json.RawMessage, error) {
method := http.MethodGet
var rdr io.Reader
if body != nil {
method = http.MethodPost
b, _ := json.Marshal(body)
rdr = bytes.NewReader(b)
}
req, _ := http.NewRequest(method, base+"/"+path, rdr)
req.Header.Set("Authorization", "Bearer "+token)
if body != nil {
req.Header.Set("Content-Type", "application/json")
}
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
var env envelope
if err := json.NewDecoder(res.Body).Decode(&env); err != nil {
return nil, err
}
if !env.OK {
return nil, fmt.Errorf("%s: %s", env.Error.Code, env.Error.Message)
}
return env.Data, nil
}
import java.net.URI;
import java.net.http.*;
public class IvDesk {
static final String BASE = "https://api.skillsafe.ai/v1/app-api";
static final String SLUG = "iv-desk";
static final String TOKEN = System.getenv().getOrDefault("SKILLSAFE_TOKEN", "YOUR_TOKEN");
static final HttpClient HTTP = HttpClient.newHttpClient();
static String call(String path, String jsonBody) throws Exception {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create(BASE + "/" + path))
.header("Authorization", "Bearer " + TOKEN);
if (jsonBody != null) {
b.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(jsonBody));
} else {
b.GET();
}
HttpResponse<String> res = HTTP.send(b.build(), HttpResponse.BodyHandlers.ofString());
// The envelope is always {"ok":true,"data":...} or {"ok":false,"error":...}.
return res.body();
}
}
require "json"
require "net/http"
require "uri"
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "iv-desk"
TOKEN = ENV.fetch("SKILLSAFE_TOKEN", "YOUR_TOKEN") # from https://iv-desk.skillsafe.ai/tokens.html
def call(path, body = nil)
uri = URI("#{BASE}/#{path}")
req = body ? Net::HTTP::Post.new(uri) : Net::HTTP::Get.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
if body
req["Content-Type"] = "application/json"
req.body = JSON.generate(body)
end
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
payload = JSON.parse(res.body)
raise "#{payload['error']['code']}: #{payload['error']['message']}" unless payload["ok"]
payload["data"]
end
<?php
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "iv-desk";
define("TOKEN", getenv("SKILLSAFE_TOKEN") ?: "YOUR_TOKEN"); // from /tokens.html
function call(string $path, ?array $body = null) {
$ch = curl_init(BASE . "/" . $path);
$headers = ["Authorization: Bearer " . TOKEN];
if ($body !== null) {
$headers[] = "Content-Type: application/json";
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($body));
}
curl_setopt($ch, CURLOPT_HTTPHEADER, $headers);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$payload = json_decode(curl_exec($ch), true);
curl_close($ch);
if (empty($payload["ok"])) {
throw new RuntimeException($payload["error"]["code"] . ": " . $payload["error"]["message"]);
}
return $payload["data"];
}
using System.Net.Http.Json;
using System.Text.Json;
static class IvDesk
{
const string Base = "https://api.skillsafe.ai/v1/app-api";
const string Slug = "iv-desk";
static readonly string Token =
Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN";
static readonly HttpClient Http = new();
public static async Task<JsonElement> Call(string path, object? body = null)
{
var req = new HttpRequestMessage(body is null ? HttpMethod.Get : HttpMethod.Post, $"{Base}/{path}");
req.Headers.Add("Authorization", $"Bearer {Token}");
if (body is not null) req.Content = JsonContent.Create(body);
var res = await Http.SendAsync(req);
var payload = await res.Content.ReadFromJsonAsync<JsonElement>();
if (!payload.GetProperty("ok").GetBoolean())
{
var e = payload.GetProperty("error");
throw new Exception($"{e.GetProperty("code")}: {e.GetProperty("message")}");
}
return payload.GetProperty("data");
}
}
2. Get a token
The easiest route is the token page: it shows the token this browser
already holds for iv-desk, with Copy token and Copy shell
export buttons, and a sign-in button for a personal token. A guest token,
minted with POST /guest and {"slug":"iv-desk"}, can call /me
and /estimate; the review is metered, so /run and /run-stream
need a personal token.
# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
# https://iv-desk.skillsafe.ai/tokens.html
# export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
-H "Content-Type: application/json" -d '{"slug":"iv-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}
# Open https://iv-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot start a metered run.
import json, urllib.request
req = urllib.request.Request(
"https://api.skillsafe.ai/v1/app-api/guest", data=b'{"slug": "iv-desk"}', method="POST")
req.add_header("Content-Type", "application/json")
with urllib.request.urlopen(req) as r:
TOKEN = json.load(r)["data"]["token"]
// Open https://iv-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
const res = await fetch("https://api.skillsafe.ai/v1/app-api/guest", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ slug: "iv-desk" }),
});
const TOKEN = (await res.json()).data.token;
// Open https://iv-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
guestReq, _ := http.NewRequest(http.MethodPost,
"https://api.skillsafe.ai/v1/app-api/guest", bytes.NewReader([]byte(`{"slug":"iv-desk"}`)))
guestReq.Header.Set("Content-Type", "application/json")
guestRes, err := http.DefaultClient.Do(guestReq)
if err != nil {
panic(err)
}
defer guestRes.Body.Close()
var guest struct {
Data struct {
Token string `json:"token"`
} `json:"data"`
}
_ = json.NewDecoder(guestRes.Body).Decode(&guest)
fmt.Println(guest.Data.Token)
// Open https://iv-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
var http = HttpClient.newHttpClient();
var guestReq = HttpRequest.newBuilder(URI.create("https://api.skillsafe.ai/v1/app-api/guest"))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString("{\"slug\":\"iv-desk\"}"))
.build();
HttpResponse<String> guest = http.send(guestReq, HttpResponse.BodyHandlers.ofString());
System.out.println(guest.body()); // {"ok":true,"data":{"token":"…","subject_type":"guest"}}
# Open https://iv-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot start a metered run.
require "json"
require "net/http"
require "uri"
uri = URI("https://api.skillsafe.ai/v1/app-api/guest")
req = Net::HTTP::Post.new(uri)
req["Content-Type"] = "application/json"
req.body = JSON.generate({ slug: "iv-desk" })
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
TOKEN = JSON.parse(res.body)["data"]["token"]
<?php
// Open https://iv-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
$ch = curl_init("https://api.skillsafe.ai/v1/app-api/guest");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode(["slug" => "iv-desk"]));
curl_setopt($ch, CURLOPT_HTTPHEADER, ["Content-Type: application/json"]);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$guest = json_decode(curl_exec($ch), true);
curl_close($ch);
echo $guest["data"]["token"];
// Open https://iv-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
using var http = new HttpClient();
var guestReq = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/guest");
guestReq.Content = new StringContent("{\"slug\":\"iv-desk\"}", Encoding.UTF8, "application/json");
var guestRes = await http.SendAsync(guestReq);
var guest = await guestRes.Content.ReadFromJsonAsync<JsonElement>();
Console.WriteLine(guest.GetProperty("data").GetProperty("token").GetString());
3. Check the session and the balance
call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}
me = call("me")
print(me["subject_type"], me.get("credits"))
const me = await call("me");
console.log(me.subject_type, me.credits);
raw, err := call("me", nil)
if err != nil {
panic(err)
}
var me struct {
SubjectType string `json:"subject_type"`
Credits int `json:"credits"`
}
_ = json.Unmarshal(raw, &me)
fmt.Println(me.SubjectType, me.Credits)
System.out.println(call("me", null));
// {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}
me = call("me")
puts "#{me['subject_type']} #{me['credits']}"
<?php
$me = call("me");
echo $me["subject_type"], " ", $me["credits"], PHP_EOL;
var me = await IvDesk.Call("me");
Console.WriteLine(me.GetProperty("subject_type").GetString());
4. Price the run (free)
/estimate returns the model binding and the credits a run would reserve. It creates no
job and charges nothing. Expect model_alias gpt-terra.
hold_credits is a reservation, not the price: it is held against your
balance while the run executes and released afterwards. min_credits is the least
balance that can start a run. What you actually pay is charged_credits, reported on the
finished job and in the done event, and it is usually far lower than the hold. The
price depends on the size of facts, so a wide file costs more than a narrow one.
The body is the input object itself, with no {"input": …} wrapper.
/estimate does not validate the body, so check the shape yourself: an object whose
every value is a string, task equal to "review", and facts
non-empty and parsing to an object.
# body.json is the input object itself - no {"input": ...} wrapper. Download it with
# the page's free "Run input .json" button, or build it with make-body.js above.
# estimate does not validate it, so check the shape first:
python3 - <<'EOF'
import json
b = json.load(open("body.json"))
assert isinstance(b, dict) and "input" not in b, "no {input: ...} wrapper"
assert all(isinstance(v, str) for v in b.values()), "every value is a string"
assert b.get("task") == "review", "the only lane is review"
assert b.get("facts", "").strip(), "facts is required"
assert isinstance(json.loads(b["facts"]), dict), "facts is a JSON string of an object"
EOF
INPUT=$(cat body.json)
call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
# "markup_bps":...,"hold_credits":...,"min_credits":...,
# "sponsor_enabled":false,"warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.
INPUT = json.load(open("body.json")) # from "Run input .json" or make-body.js
assert isinstance(INPUT, dict) and "input" not in INPUT # no wrapper
assert all(isinstance(v, str) for v in INPUT.values()) # every value a string
assert INPUT.get("task") == "review" # the only lane
assert INPUT.get("facts", "").strip() # required
assert isinstance(json.loads(INPUT["facts"]), dict) # facts is a JSON STRING
est = call("estimate", INPUT)
print(est["model_alias"], est["hold_credits"], est["min_credits"], est.get("warnings"))
me = call("me")
if me.get("credits", 0) < est["min_credits"]:
raise SystemExit("top up first: balance is below min_credits")
const INPUT = JSON.parse(readFileSync("body.json", "utf8")); // "Run input .json"
if (!INPUT || typeof INPUT !== "object" || Array.isArray(INPUT) || "input" in INPUT)
throw new Error("body must be the input object itself");
for (const [k, v] of Object.entries(INPUT))
if (typeof v !== "string") throw new Error(k + " must be a string");
if (INPUT.task !== "review") throw new Error("task must be review");
if (!(INPUT.facts || "").trim()) throw new Error("facts is required");
const facts = JSON.parse(INPUT.facts); // throws unless facts is a JSON string
if (!facts || typeof facts !== "object" || Array.isArray(facts))
throw new Error("facts must encode an object");
const est = await call("estimate", INPUT);
console.log(est.model_alias, est.hold_credits, est.min_credits, est.warnings);
const me = await call("me");
if ((me.credits ?? 0) < est.min_credits) throw new Error("top up first");
raw, _ := os.ReadFile("body.json") // from "Run input .json" or make-body.js
var input map[string]string // every field is a string, facts included
if err := json.Unmarshal(raw, &input); err != nil {
panic("body.json must be an object of strings: " + err.Error())
}
if _, wrapped := input["input"]; wrapped {
panic("send the input object itself, not {\"input\": ...}")
}
if input["task"] != "review" {
panic("task must be review")
}
if strings.TrimSpace(input["facts"]) == "" {
panic("facts is required")
}
var facts map[string]any
if err := json.Unmarshal([]byte(input["facts"]), &facts); err != nil {
panic("facts must be a JSON string holding an object")
}
est, err := call("estimate", input)
if err != nil {
panic(err)
}
fmt.Println(string(est)) // model_alias, hold_credits, min_credits
// Files, Path: java.nio.file. Use your JSON library for a strict check;
// these regexes catch the usual mistakes.
String input = Files.readString(Path.of("body.json")); // "Run input .json"
if (!input.matches("(?s)\\s*\\{.*\\}\\s*"))
throw new IllegalStateException("body.json must be one JSON object");
if (input.matches("(?s)\\s*\\{\\s*\"input\"\\s*:.*"))
throw new IllegalStateException("no {\"input\": ...} wrapper");
if (!input.matches("(?s).*\"task\"\\s*:\\s*\"review\".*"))
throw new IllegalStateException("task must be review");
// facts is a STRING: its value starts with a quote, then an escaped object.
if (!input.matches("(?s).*\"facts\"\\s*:\\s*\"\\{.*"))
throw new IllegalStateException("facts must be a JSON-encoded string");
String est = call("estimate", input);
System.out.println(est); // model_alias, hold_credits, min_credits
INPUT = JSON.parse(File.read("body.json")) # "Run input .json"
raise "body must be an object" unless INPUT.is_a?(Hash) && !INPUT.key?("input")
INPUT.each { |k, v| raise "#{k} must be a string" unless v.is_a?(String) }
raise "task must be review" unless INPUT["task"] == "review"
raise "facts is required" if INPUT["facts"].to_s.strip.empty?
raise "facts must encode an object" unless JSON.parse(INPUT["facts"]).is_a?(Hash)
est = call("estimate", INPUT)
puts est["model_alias"], est["hold_credits"], est["min_credits"]
<?php
$input = json_decode(file_get_contents("body.json"), true); // "Run input .json"
if (!is_array($input) || array_is_list($input) || isset($input["input"])) {
throw new Exception("body must be the input object itself");
}
foreach ($input as $k => $v) {
if (!is_string($v)) { throw new Exception("$k must be a string"); }
}
if (($input["task"] ?? "") !== "review") { throw new Exception("task must be review"); }
if (trim($input["facts"] ?? "") === "") { throw new Exception("facts is required"); }
$facts = json_decode($input["facts"], true);
if (!is_array($facts) || array_is_list($facts)) {
throw new Exception("facts must be a JSON string of an object");
}
$est = call("estimate", $input);
echo $est["model_alias"], " ", $est["hold_credits"], " ", $est["min_credits"], PHP_EOL;
var input = File.ReadAllText("body.json"); // "Run input .json"
using var doc = JsonDocument.Parse(input);
var root = doc.RootElement;
if (root.ValueKind != JsonValueKind.Object || root.TryGetProperty("input", out _))
throw new Exception("body must be the input object itself");
foreach (var p in root.EnumerateObject())
if (p.Value.ValueKind != JsonValueKind.String)
throw new Exception($"{p.Name} must be a string");
if (root.GetProperty("task").GetString() != "review")
throw new Exception("task must be review");
var factsText = root.GetProperty("facts").GetString() ?? "";
if (factsText.Trim() == "") throw new Exception("facts is required");
using var facts = JsonDocument.Parse(factsText); // facts is a JSON string
if (facts.RootElement.ValueKind != JsonValueKind.Object)
throw new Exception("facts must encode an object");
var est = await IvDesk.Call("estimate", root);
Console.WriteLine(est); // model_alias, hold_credits, min_credits
5. Run it, then poll
POST /run returns a job_id; poll GET /jobs/{id} until it is
terminal. The reply is a string at data.output.output: JSON.parse
it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the
attempt number, iv-desk:review:<hash>:a<attempt> (for example
iv-desk:review:8bcf46894487e343:a1), so a retried request returns the same job instead
of billing a second run. Use one key per distinct input: a rescreened file, changed settings or a
new question is a new hash, and replaying an old key with a different body is a 409.
The page uses IvKit.hashInput(body) for the hash (make-body.js prints that
key); any stable digest of the body works from other languages. Leave retry_note out of
the hash and bump the attempt instead.
# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
# One lane, so the key is iv-desk:review:<hash>:a<attempt>.
KEY="iv-desk:review:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "$BASE/run" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')
while :; do
OUT=$(call "jobs/$JOB")
STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
[ "$STATUS" = "succeeded" ] && break
[ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
sleep 2
done
# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
# "output":{"output":"{\"task\":\"review\",\"verdict\":\"revise\", ...}"},
# "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c \
'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json
import hashlib, time
digest = hashlib.sha256(json.dumps(INPUT, sort_keys=True).encode()).hexdigest()[:16]
key = f"iv-desk:review:{digest}:a1"
req = urllib.request.Request(f"{BASE}/run", data=json.dumps(INPUT).encode(), method="POST")
req.add_header("Authorization", f"Bearer {TOKEN}")
req.add_header("Content-Type", "application/json")
req.add_header("Idempotency-Key", key)
with urllib.request.urlopen(req) as r:
job_id = json.load(r)["data"]["job_id"]
while True:
job = call(f"jobs/{job_id}")
if job["status"] in ("succeeded", "failed"):
break
time.sleep(2)
if job["status"] == "failed":
raise RuntimeError(job.get("error"))
text = job["output"]["output"] # the reply, as a string
print("charged", job.get("charged_credits"), "truncated", job.get("truncated"))
import { createHash } from "node:crypto";
const digest = createHash("sha256").update(JSON.stringify(INPUT)).digest("hex").slice(0, 16);
const key = `iv-desk:review:${digest}:a1`;
const started = await fetch(`${BASE}/run`, {
method: "POST",
headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json", "Idempotency-Key": key },
body: JSON.stringify(INPUT),
}).then((r) => r.json());
if (!started.ok) throw new Error(`${started.error.code}: ${started.error.message}`);
let job = started.data;
while (job.status !== "succeeded" && job.status !== "failed") {
await new Promise((r) => setTimeout(r, 2000));
job = await call(`jobs/${job.job_id}`);
}
if (job.status === "failed") throw new Error(JSON.stringify(job.error));
const text = job.output.output; // the reply, as a string
console.log(job.charged_credits, job.truncated);
body, _ := json.Marshal(input)
sum := sha256.Sum256(body)
key := fmt.Sprintf("iv-desk:review:%x:a1", sum[:8])
req, _ := http.NewRequest(http.MethodPost, base+"/run", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
var started struct {
Data struct {
JobID string `json:"job_id"`
} `json:"data"`
}
_ = json.NewDecoder(res.Body).Decode(&started)
res.Body.Close()
var jobOutput string
for {
raw, err := call("jobs/"+started.Data.JobID, nil)
if err != nil {
panic(err)
}
var job struct {
Status string `json:"status"`
Output struct {
Output string `json:"output"`
} `json:"output"`
Charged int `json:"charged_credits"`
Truncated bool `json:"truncated"`
}
_ = json.Unmarshal(raw, &job)
if job.Status == "succeeded" {
jobOutput = job.Output.Output
fmt.Println(job.Charged, job.Truncated)
break
}
if job.Status == "failed" {
panic(string(raw))
}
time.Sleep(2 * time.Second)
}
String key = "iv-desk:review:" + sha256Hex(input).substring(0, 16) + ":a1";
HttpRequest run = HttpRequest.newBuilder(URI.create(BASE + "/run"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.POST(HttpRequest.BodyPublishers.ofString(input)).build();
String started = HTTP.send(run, HttpResponse.BodyHandlers.ofString()).body();
String jobId = started.replaceAll(".*\"job_id\":\"([^\"]+)\".*", "$1");
while (true) {
String job = call("jobs/" + jobId, null);
if (job.contains("\"status\":\"succeeded\"")) { System.out.println(job); break; }
if (job.contains("\"status\":\"failed\"")) throw new RuntimeException(job);
Thread.sleep(2000);
}
// Parse data.output.output (a string holding the reply JSON) with your JSON library.
// sha256Hex: HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(input.getBytes(UTF_8)))
require "digest"
key = "iv-desk:review:#{Digest::SHA256.hexdigest(JSON.generate(INPUT))[0, 16]}:a1"
uri = URI("#{BASE}/run")
req = Net::HTTP::Post.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
req["Content-Type"] = "application/json"
req["Idempotency-Key"] = key
req.body = JSON.generate(INPUT)
job = JSON.parse(Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }.body)["data"]
until %w[succeeded failed].include?(job["status"])
sleep 2
job = call("jobs/#{job['job_id']}")
end
raise job.inspect if job["status"] == "failed"
text = job["output"]["output"] # the reply, as a string
puts job["charged_credits"], job["truncated"]
<?php
$key = "iv-desk:review:" . substr(hash("sha256", json_encode($input)), 0, 16) . ":a1";
$ch = curl_init(BASE . "/run");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => json_encode($input),
CURLOPT_HTTPHEADER => [
"Authorization: Bearer " . TOKEN,
"Content-Type: application/json",
"Idempotency-Key: " . $key,
],
CURLOPT_RETURNTRANSFER => true,
]);
$job = json_decode(curl_exec($ch), true)["data"];
curl_close($ch);
while (!in_array($job["status"], ["succeeded", "failed"], true)) {
sleep(2);
$job = call("jobs/" . $job["job_id"]);
}
if ($job["status"] === "failed") { throw new RuntimeException(json_encode($job)); }
$text = $job["output"]["output"]; // the reply, as a string
echo $job["charged_credits"], PHP_EOL;
using System.Security.Cryptography;
var json = input; // the body.json text from step 4
var digest = Convert.ToHexString(SHA256.HashData(System.Text.Encoding.UTF8.GetBytes(json)));
var key = "iv-desk:review:" + digest[..16].ToLower() + ":a1";
var req = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/run");
req.Headers.Add("Authorization", $"Bearer {Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN"}");
req.Headers.Add("Idempotency-Key", key);
req.Content = new StringContent(json, System.Text.Encoding.UTF8, "application/json");
var started = await (await new HttpClient().SendAsync(req)).Content.ReadFromJsonAsync<JsonElement>();
var jobId = started.GetProperty("data").GetProperty("job_id").GetString();
JsonElement job;
while (true)
{
job = await IvDesk.Call($"jobs/{jobId}");
var status = job.GetProperty("status").GetString();
if (status == "succeeded") break;
if (status == "failed") throw new Exception(job.ToString());
await Task.Delay(2000);
}
var output = job.GetProperty("output").GetProperty("output").GetString()!; // the reply, as a string
6. Or stream it
POST /run-stream takes the same body and headers and answers with server-sent events:
job (the job id), delta (chunks of the reply) and done (the
status, charged_credits, truncated and, when present, the full
output). A browser page may receive only tick heartbeats and then
done, never a delta, so take the reply from done.output.output
when it is there, fall back to the concatenated deltas, and fall back again to
GET /jobs/{id}.
# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-H "Accept: text/event-stream" \
-d "$INPUT"
# event: job {"job_id":"job_..."}
# event: delta {"text":"{\"task\":\"review\",\"verdict\":\"revise\",\"headline\":\"This"}
# event: done {"status":"succeeded","charged_credits":...,"truncated":false}
req = urllib.request.Request(f"{BASE}/run-stream", data=json.dumps(INPUT).encode(), method="POST")
for h, v in (("Authorization", f"Bearer {TOKEN}"), ("Content-Type", "application/json"),
("Idempotency-Key", key), ("Accept", "text/event-stream")):
req.add_header(h, v)
raw, done, event = "", {}, None
with urllib.request.urlopen(req) as stream:
for line in stream:
line = line.decode().rstrip("\n")
if line.startswith("event: "):
event = line[7:]
elif line.startswith("data: ") and event == "delta":
raw += json.loads(line[6:]).get("text", "")
elif line.startswith("data: ") and event == "done":
done = json.loads(line[6:])
text = (done.get("output") or {}).get("output") or raw
print(done.get("status"), done.get("charged_credits"), done.get("truncated"))
const res = await fetch(`${BASE}/run-stream`, {
method: "POST",
headers: {
Authorization: `Bearer ${TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": key,
Accept: "text/event-stream",
},
body: JSON.stringify(INPUT),
});
const reader = res.body.getReader();
const dec = new TextDecoder();
let buf = "", raw = "", event = null, done = null;
for (;;) {
const { value, done: end } = await reader.read();
if (end) break;
buf += dec.decode(value, { stream: true });
let i;
while ((i = buf.indexOf("\n")) >= 0) {
const line = buf.slice(0, i); buf = buf.slice(i + 1);
if (line.startsWith("event: ")) event = line.slice(7);
else if (line.startsWith("data: ") && event === "delta") raw += JSON.parse(line.slice(6)).text || "";
else if (line.startsWith("data: ") && event === "done") done = JSON.parse(line.slice(6));
}
}
const streamed = done?.output?.output || raw; // browsers may get only ticks + done
console.log(done, streamed.length);
req, _ = http.NewRequest(http.MethodPost, base+"/run-stream", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
req.Header.Set("Accept", "text/event-stream")
res, err = http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
var raw strings.Builder
event := ""
sc := bufio.NewScanner(res.Body)
sc.Buffer(make([]byte, 1<<20), 1<<20)
for sc.Scan() {
line := sc.Text()
switch {
case strings.HasPrefix(line, "event: "):
event = line[7:]
case strings.HasPrefix(line, "data: ") && event == "delta":
var d struct{ Text string `json:"text"` }
_ = json.Unmarshal([]byte(line[6:]), &d)
raw.WriteString(d.Text)
case strings.HasPrefix(line, "data: ") && event == "done":
fmt.Println("done:", line[6:])
}
}
HttpRequest stream = HttpRequest.newBuilder(URI.create(BASE + "/run-stream"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.header("Accept", "text/event-stream")
.POST(HttpRequest.BodyPublishers.ofString(input)).build();
HTTP.send(stream, HttpResponse.BodyHandlers.ofLines()).body().forEach(line -> {
// "event: delta" lines are followed by "data: {\"text\":...}"; "event: done" by the status.
if (line.startsWith("data: ")) System.out.println(line.substring(6));
});
uri = URI("#{BASE}/run-stream")
req = Net::HTTP::Post.new(uri)
{ "Authorization" => "Bearer #{TOKEN}", "Content-Type" => "application/json",
"Idempotency-Key" => key, "Accept" => "text/event-stream" }.each { |k, v| req[k] = v }
req.body = JSON.generate(INPUT)
raw, event = +"", nil
Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) do |h|
h.request(req) do |res|
res.read_body do |chunk|
chunk.each_line do |line|
line = line.chomp
if line.start_with?("event: ") then event = line[7..]
elsif line.start_with?("data: ") && event == "delta" then raw << JSON.parse(line[6..])["text"].to_s
elsif line.start_with?("data: ") && event == "done" then puts line[6..]
end
end
end
end
end
<?php
$raw = ""; $event = null;
$ch = curl_init(BASE . "/run-stream");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => json_encode($input),
CURLOPT_HTTPHEADER => [
"Authorization: Bearer " . TOKEN,
"Content-Type: application/json",
"Idempotency-Key: " . $key,
"Accept: text/event-stream",
],
CURLOPT_WRITEFUNCTION => function ($ch, $chunk) use (&$raw, &$event) {
foreach (explode("\n", $chunk) as $line) {
if (str_starts_with($line, "event: ")) $event = substr($line, 7);
elseif (str_starts_with($line, "data: ") && $event === "delta") $raw .= json_decode(substr($line, 6), true)["text"] ?? "";
elseif (str_starts_with($line, "data: ") && $event === "done") echo substr($line, 6), PHP_EOL;
}
return strlen($chunk);
},
]);
curl_exec($ch);
curl_close($ch);
var sreq = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/run-stream");
sreq.Headers.Add("Authorization", $"Bearer {Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN"}");
sreq.Headers.Add("Idempotency-Key", key);
sreq.Headers.Add("Accept", "text/event-stream");
sreq.Content = new StringContent(json, System.Text.Encoding.UTF8, "application/json");
using var sres = await new HttpClient().SendAsync(sreq, HttpCompletionOption.ResponseHeadersRead);
using var sr = new StreamReader(await sres.Content.ReadAsStreamAsync());
var raw = new System.Text.StringBuilder(); string? ev = null, line;
while ((line = await sr.ReadLineAsync()) != null)
{
if (line.StartsWith("event: ")) ev = line[7..];
else if (line.StartsWith("data: ") && ev == "delta") raw.Append(JsonSerializer.Deserialize<JsonElement>(line[6..]).GetProperty("text").GetString());
else if (line.StartsWith("data: ") && ev == "done") Console.WriteLine(line[6..]);
}
7. Parse the reply
The reply is one JSON object, delivered as a string in data.output.output; you must
JSON.parse it. The model is told to send no code fences, but tolerate them: strip a
leading ```json and a trailing ```, keep everything from the first
{ to the last }, and parse that outer object. The keys are the same on
every run; see the output contract.
# reply.json holds data.output.output from step 5. Strip any fence, keep the object:
python3 - <<'EOF'
import json, re
t = open("reply.json").read().strip()
t = re.sub(r"^```(?:json)?\s*", "", t, flags=re.I)
t = re.sub(r"\s*```\s*$", "", t)
r = json.loads(t[t.index("{"):t.rindex("}") + 1])
print(r["task"], r["verdict"], "-", r["headline"])
for s in r["step_reviews"]:
print("STEP", s["step"], s["verdict"], "|", s["note"])
for c in r["feature_calls"]:
print("CALL", c["feature"], c["call"], "(screen:", c["engine"], "iv", c["iv"], ")")
for s in r["leakage_suspects"]:
print("LEAK?", s["feature"], "| check:", s["check"])
for a in r["threshold_advice"]:
print("SET", a["parameter"], a["current"], "->", a["suggested"])
EOF
import re
t = re.sub(r"^```(?:json)?\s*", "", text.strip(), flags=re.I)
t = re.sub(r"\s*```\s*$", "", t)
reply = json.loads(t[t.index("{"):t.rindex("}") + 1])
print(reply["verdict"], reply["headline"])
print([(s["step"], s["verdict"]) for s in reply["step_reviews"]])
print([(c["feature"], c["call"], c["engine"]) for c in reply["feature_calls"]])
print([s["feature"] for s in reply["leakage_suspects"]])
print([(a["parameter"], a["current"], a["suggested"]) for a in reply["threshold_advice"]])
let t = text.trim().replace(/^```(?:json)?\s*/i, "").replace(/\s*```\s*$/, "");
const reply = JSON.parse(t.slice(t.indexOf("{"), t.lastIndexOf("}") + 1));
console.log(reply.verdict, reply.headline);
console.log(reply.step_reviews.map((s) => `${s.step}: ${s.verdict}`));
console.log(reply.feature_calls.map((c) => `${c.feature} ${c.call} (${c.engine})`));
console.log(reply.leakage_suspects.map((s) => `${s.feature} -> ${s.check}`));
console.log(reply.threshold_advice.map((a) => `${a.parameter} ${a.current} -> ${a.suggested}`));
t := strings.TrimSpace(jobOutput)
t = strings.TrimPrefix(strings.TrimPrefix(t, "```json"), "```")
t = strings.TrimSuffix(strings.TrimSpace(t), "```")
start, end := strings.Index(t, "{"), strings.LastIndex(t, "}")
var reply struct {
Task string `json:"task"`
Verdict string `json:"verdict"`
Headline string `json:"headline"`
TLDR []string `json:"tldr"`
StepReviews []struct {
Step, Verdict, Note string
} `json:"step_reviews"`
FeatureCalls []struct {
Feature, Call, Engine, Reason, Ref string
IV *float64 `json:"iv"` // null allowed
} `json:"feature_calls"`
LeakageSuspects []struct {
Feature, Evidence, Check string
} `json:"leakage_suspects"`
ThresholdAdvice []struct {
Parameter string
Current any // a number, or a string for sentinels
Suggested any
Why string
} `json:"threshold_advice"`
NextSteps []string `json:"next_steps"`
Memo []string `json:"memo"`
PrescanResponses []struct {
Ref, Verdict, Note string
} `json:"prescan_responses"`
}
if err := json.Unmarshal([]byte(t[start:end+1]), &reply); err != nil {
panic(err)
}
fmt.Println(reply.Verdict, reply.Headline)
for _, c := range reply.FeatureCalls {
fmt.Println(c.Feature, c.Call, c.Engine)
}
// output is data.output.output from step 5: a string holding the reply JSON.
String t = output.strip().replaceFirst("^```(?:json)?\\s*", "")
.replaceFirst("\\s*```\\s*$", "");
String json = t.substring(t.indexOf('{'), t.lastIndexOf('}') + 1);
// With Jackson: JsonNode r = new ObjectMapper().readTree(json);
// r.get("verdict"): ready | revise | blocked
// r.get("step_reviews"): 7 entries, sample months missing iv psi noise corr
// r.get("feature_calls"): {feature, call, engine, iv, reason, ref}
// r.get("leakage_suspects"), r.get("threshold_advice"), r.get("memo")
System.out.println(json);
t = text.strip.sub(/\A```(?:json)?\s*/i, "").sub(/\s*```\s*\z/, "")
reply = JSON.parse(t[t.index("{")..t.rindex("}")])
puts reply["verdict"], reply["headline"]
reply["step_reviews"].each { |s| puts "#{s['step']}: #{s['verdict']}" }
reply["feature_calls"].each { |c| puts "#{c['feature']} #{c['call']} (#{c['engine']})" }
reply["leakage_suspects"].each { |s| puts "#{s['feature']} check: #{s['check']}" }
reply["threshold_advice"].each do |a|
puts "#{a['parameter']} #{a['current']} -> #{a['suggested']}"
end
<?php
$t = preg_replace('/\s*```\s*$/', "", preg_replace('/^```(?:json)?\s*/i', "", trim($text)));
$a = strpos($t, "{");
$reply = json_decode(substr($t, $a, strrpos($t, "}") - $a + 1), true);
echo $reply["verdict"], " ", $reply["headline"], PHP_EOL;
foreach ($reply["step_reviews"] as $s) { echo $s["step"], ": ", $s["verdict"], PHP_EOL; }
foreach ($reply["feature_calls"] as $c) {
echo $c["feature"], " ", $c["call"], " (", $c["engine"], ")", PHP_EOL;
}
foreach ($reply["leakage_suspects"] as $s) { echo $s["feature"], " check: ", $s["check"], PHP_EOL; }
var t = System.Text.RegularExpressions.Regex.Replace(output.Trim(), @"^```(?:json)?\s*", "");
t = System.Text.RegularExpressions.Regex.Replace(t, @"\s*```\s*$", "");
var json2 = t.Substring(t.IndexOf('{'), t.LastIndexOf('}') - t.IndexOf('{') + 1);
using var parsed = JsonDocument.Parse(json2);
var r = parsed.RootElement;
Console.WriteLine($"{r.GetProperty("verdict")} {r.GetProperty("headline")}");
foreach (var s in r.GetProperty("step_reviews").EnumerateArray())
Console.WriteLine($"{s.GetProperty("step")}: {s.GetProperty("verdict")}");
foreach (var c in r.GetProperty("feature_calls").EnumerateArray())
Console.WriteLine($"{c.GetProperty("feature")} {c.GetProperty("call")} ({c.GetProperty("engine")})");
Invariants worth asserting
The web page holds every reply to the screen before it shows it (recon.js,
reconcile()) and lists any disagreement next to the review. Do the same before you file
a memo:
- It parses: the reply is one JSON object (after stripping any fence), all eleven keys are present, and
taskis"review". - Flags: every flag in
facts.flagshas exactly oneprescan_responsesentry; no response names a flag that was not sent; a dismissal says which fact shows the flag is harmless (read it and decide whether you agree). - Nothing invented: every feature named in
feature_callsandleakage_suspectsis infacts.features,facts.survivorsor a flag'sfeatures, spelled exactly. - The screen quoted exactly:
engineiskeptwhen the feature'sdropsis"kept", elsedropped;ivequals the feature'sivinfacts(within 0.00015), or isnull. - Survivors covered: every survivor (the first 30) has a call.
- Leakage: a feature in
leakage_suspectsis never calledkeep; every feature of a confirmedhighormediumleakageflag is a suspect, or calleddroporinvestigate, or is kept with its name in that flag's note (the model cleared it, as the example does forbureau_score). - Settings: every
threshold_adviceentry names a key offacts.settingsand quotes its current value. - Steps: exactly seven
step_reviews, one per step. - Verdict floor: never looser than the flags left standing (those not dismissed): a
highsampleflag, or no survivors, meansblocked; any otherhighormediummeans at leastrevise. - Overrides are not errors: a
keepon a feature the screen dropped is allowed, but itsreasonmust name the step and say why the rule is wrong for this sample. List them for a human.
# assert-reply.py - the page's reconciliation checks (recon.js), for body.json
# and reply.json from the steps above.
import json
body = json.load(open("body.json")); facts = json.loads(body["facts"])
t = open("reply.json").read(); r = json.loads(t[t.index("{"):t.rindex("}") + 1])
KEYS = ["task", "verdict", "headline", "tldr", "step_reviews", "feature_calls",
"leakage_suspects", "threshold_advice", "next_steps", "memo", "prescan_responses"]
STEPS = ["sample", "months", "missing", "iv", "psi", "noise", "corr"]
missing = [k for k in KEYS if k not in r]
assert not missing, "keys missing: " + ", ".join(missing)
assert r["task"] == "review", "task is " + str(r["task"])
assert r["verdict"] in ("ready", "revise", "blocked")
# Steps: exactly one review per step, in order.
assert [s["step"] for s in r["step_reviews"]] == STEPS, "one review per step, in order"
assert all(s["verdict"] in ("agree", "adjust", "question") for s in r["step_reviews"])
# Flags: each answered exactly once, none invented.
flags = {f["id"]: f for f in facts["flags"]}
answered = [p["ref"].upper() for p in r["prescan_responses"]]
assert sorted(answered) == sorted(flags), "every flag answered exactly once"
assert all(p["verdict"] in ("confirmed", "dismissed") for p in r["prescan_responses"])
resp = {p["ref"].upper(): p for p in r["prescan_responses"]}
# Features: real names, the screen's keep/drop and IV quoted exactly.
feats = {f["name"]: f for f in facts["features"]}
known = set(feats) | set(facts["survivors"]) | {n for f in facts["flags"] for n in f["features"]}
calls = {c["feature"]: c for c in r["feature_calls"]}
named = list(calls) + [s["feature"] for s in r["leakage_suspects"]]
assert all(n in known for n in named), "a feature not in your file"
for c in r["feature_calls"]:
assert c["call"] in ("keep", "drop", "investigate")
f = feats.get(c["feature"])
if not f:
continue # named in a flag but not detailed
assert c["engine"] == ("kept" if f["drops"] == "kept" else "dropped"), c["feature"]
if c["iv"] is not None:
assert "iv" in f and abs(c["iv"] - f["iv"]) <= 0.00015, "IV misquoted: " + c["feature"]
# Survivors covered; a leakage suspect is never kept.
assert all(n in calls for n in facts["survivors"][:30]), "a survivor has no call"
suspects = {s["feature"] for s in r["leakage_suspects"]}
assert not [n for n in suspects if calls.get(n, {}).get("call") == "keep"], "suspect kept"
# Confirmed high/medium leakage flags carried (or explicitly cleared in the note).
for fid, f in flags.items():
p = resp[fid]
if f["category"] != "leakage" or f["severity"] == "low" or p["verdict"] != "confirmed":
continue
for n in f["features"]:
kept = calls.get(n, {}).get("call") == "keep"
assert n in suspects or (n in calls and not kept) or (kept and n in p["note"]), n
# Threshold advice names a real setting and quotes its current value.
S = facts["settings"]
for a in r["threshold_advice"]:
assert a["parameter"] in S, "no such setting: " + a["parameter"]
cur, real = a["current"], S[a["parameter"]]
if isinstance(real, str):
assert str(cur).replace(" ", "") == real.replace(" ", ""), a["parameter"]
else:
assert abs(float(cur) - real) <= max(1e-6, abs(real) * 0.002), a["parameter"]
# Verdict: never looser than the flags left standing.
rank = {"ready": 0, "revise": 1, "blocked": 2}
floor = 0
for fid, f in flags.items():
if resp[fid]["verdict"] == "dismissed":
continue
lvl = 2 if f["severity"] == "high" and f["category"] == "sample" else 1 if f["severity"] != "low" else 0
floor = max(floor, lvl)
if not facts["survivors"]:
floor = 2
assert rank[r["verdict"]] >= floor, "verdict looser than the flags left standing"
overrides = [c["feature"] for c in r["feature_calls"] if c["call"] == "keep" and c["engine"] == "dropped"]
print("ok:", r["verdict"], len(r["feature_calls"]), "calls;", "overrides:", overrides or "none")
The worked example passes every one of these checks, and the page's own
reconcile() reports zero disagreements for it.
The output contract
The reply is one JSON object with eleven keys, always all present, in this order. Arrays may be
empty ([], never a filler such as "None"). An enum is written
"a|b|c": the reply carries exactly one of the values. Text fields are plain prose, no
Markdown, no emoji, each under 400 characters except memo paragraphs (up to 900). Every
feature name and every number comes from facts.
{"task":"review",
"verdict":"ready|revise|blocked",
"headline":"...",
"tldr":["...","..."],
"step_reviews":[{"step":"sample","verdict":"agree|adjust|question","note":"..."}],
"feature_calls":[{"feature":"...","call":"keep|drop|investigate",
"engine":"kept|dropped","iv":0.1234,"reason":"...","ref":"F2"}],
"leakage_suspects":[{"feature":"...","evidence":"...","check":"..."}],
"threshold_advice":[{"parameter":"psi_max_orgs","current":6,"suggested":2,"why":"..."}],
"next_steps":["...","..."],
"memo":["...","..."],
"prescan_responses":[{"ref":"F1","verdict":"confirmed|dismissed","note":"..."}]}
| key | shape | what it holds |
|---|---|---|
task | string | Always "review". |
verdict | enum | Can this screen go to binning and modeling; see the next table. |
headline | string | One sentence saying whether the screen can go to modeling. |
tldr | array of strings | 2-5 bullets. When you sent a question, one bullet starts "Answer:". |
step_reviews | array of {step, verdict, note} | Exactly seven, in order: sample, months, missing, iv, psi, noise, corr. agree: the rule and its result fit the sample. adjust: a threshold should change (given in threshold_advice). question: the result cannot be trusted as run, for example a rule that could not fire as set. |
feature_calls | array of {feature, call, engine, iv, reason, ref} | At most 40: every survivor (up to 30), every feature of a confirmed flag, and any dropped feature worth bringing back. engine is what the screen decided; iv is copied from facts or null; ref is the flag id the call answers, or "". A keep on an engine dropped feature is an override and its reason names the step. |
leakage_suspects | array of {feature, evidence, check} | Features that may encode the outcome, with the evidence from facts and the concrete check that would clear or convict each one. May include a feature the browser did not flag, saying so in evidence. |
threshold_advice | array of {parameter, current, suggested, why} | At most 6. parameter is a key of facts.settings, current its value there; suggested is a number, or a comma-separated list for sentinels. |
next_steps | array of strings | 3-7 concrete actions, in order. |
memo | array of strings | 2-4 paragraphs a model validator could file as the cleaning report: the sample, what each step removed, the leakage and stability concerns, the survivors and what happens next. |
prescan_responses | array of {ref, verdict, note} | Exactly one per flag in facts.flags: confirmed (a real issue for this sample) or dismissed (the note names the fact that shows it is harmless). Empty when no flags were sent. |
Verdict
| verdict | meaning |
|---|---|
blocked | The sample cannot support a variable screen yet: a confirmed high flag of category sample (too few bads, nothing screened, nothing survives), or no survivors. |
revise | Usable, but something must change before modeling: any other confirmed high flag (leakage) or any confirmed medium flag. |
ready | Nothing confirmed at high or medium; the survivors can go to binning and modeling. |
The verdict floor: the model may be stricter than
facts.browser_verdict, never looser, unless it dismissed the flags that set it. Your
own check computes the floor from the flags left standing, exactly as in the
invariants.
Enums
| where | values | the page's fallback |
|---|---|---|
verdict | ready, revise, blocked | revise |
step_reviews[].step | sample, months, missing, iv, psi, noise, corr (in this order) | kept as written |
step_reviews[].verdict | agree, adjust, question | question |
feature_calls[].call | keep, drop, investigate | investigate |
feature_calls[].engine | kept, dropped | "" |
prescan_responses[].verdict | confirmed, dismissed | confirmed |
Worked example: review
The page's Retail instalment loans example: 22,273 loan rows from three lenders
(bank_a, bank_b, broker_c) plus partner_d, held
out of sample; 19,374 modeling rows with 2,122 bads over 12 months. Twenty features were screened and
6 survive. The screen raised 6 flags: F2 high leakage (overdue_days_m3 has
IV 16.2826, bureau_score 0.5756), F1 and F5 medium PSI (the
rule needs 6 unstable organizations and there are 3, so fraud_score_v2, unstable in all
three, was not dropped by PSI), F3 medium missing (sentinels in
features that hold other negatives), F6 medium correlation (a pair with r 0.9998 left
in place by top_n_keep) and F4 low (a constant column), so
browser_verdict is revise. Its Idempotency-Key from the page is
iv-desk:review:8bcf46894487e343:a1.
The request body. The facts string is 6,892 characters; only its first 150 are shown here,
ending in a marked …. Every other field is shown in full, exactly as sent:
{
"task": "review",
"dataset": "retail-instalment-2025",
"context": "Unsecured instalment loans, three lenders plus one partner held out of sample. Target = 1 if 60+ days past due within 12 months of booking. Features are meant to be known at application time.",
"facts": "{\"roles\":{\"date\":\"apply_date\",\"label\":\"target\",\"org\":\"org_info\",\"keys\":[\"loan_id\"]},\"settings\":{\"sentinels\":\"-1, -999, -1111\",\"min_ym_bad\":10,\"min_ym_…",
"question": "Which of these would you take into a logistic scorecard?"
}
The same facts, decoded. Real values throughout; the lines starting with
… are trims made for this page, not part of the JSON:
{
"roles": {"date": "apply_date", "label": "target", "org": "org_info", "keys": ["loan_id"]},
"settings": {
"sentinels": "-1, -999, -1111", "min_ym_bad": 10, "min_ym_total": 500,
"missing_ratio": 0.6, "overall_iv": 0.1, "org_iv": 0.1,
"max_low_iv_orgs": 2, "psi": 0.1, "psi_months_ratio": 0.3333,
"psi_max_orgs": 6, "psi_min_month": 100, "null_perms": 10,
"max_corr": 0.9, "top_n_keep": 20
},
"sample": {
"rows_in": 22273, "invalid_label": 0, "invalid_date": 0,
"duplicates": 0, "rows_valid": 22273, "oos_rows": 2899,
"modeling_rows_before_months": 19374, "modeling_rows": 19374, "modeling_bads": 2122,
"modeling_bad_rate": 0.1095, "months_total": 12, "months_dropped": 0
},
"organizations": [
{"org": "bank_a", "type": "modeling", "rows": 7592, "bads": 818, "bad_rate": 0.1077, "months": 12},
{"org": "bank_b", "type": "modeling", "rows": 6647, "bads": 655, "bad_rate": 0.0985, "months": 12},
{"org": "broker_c", "type": "modeling", "rows": 5135, "bads": 649, "bad_rate": 0.1264, "months": 12},
{"org": "partner_d", "type": "oos", "rows": 2899, "bads": 329, "bad_rate": 0.1135, "months": 12}
],
"months": [
{"ym": "202501", "rows": 1643, "bads": 191, "dropped": false, "reason": ""},
{"ym": "202502", "rows": 1591, "bads": 176, "dropped": false, "reason": ""}
… 10 more months, 202503 to 202512, all with dropped false
],
"step_counts": {"constant": 1, "missing": 1, "iv": 13, "psi": 0, "noise": 8, "corr": 0},
"features_total": 20,
"features_detailed": 20,
"features": [
{"name": "overdue_days_m3", "type": "numeric", "missing": 0, "iv": 16.2826, "drops": "kept", "psi_max": 0.0219, "null_margin": 16.2709},
{"name": "bureau_score", "type": "numeric", "missing": 0, "iv": 0.5756, "drops": "kept", "psi_max": 0.11, "null_margin": 0.5633},
{"name": "util_revolving", "type": "numeric", "missing": 0, "iv": 0.14, "drops": "kept", "psi_max": 0.1154, "null_margin": 0.1259, "top_corr": "util_revolving_pct 0.9998"},
{"name": "util_revolving_pct", "type": "numeric", "missing": 0, "iv": 0.1393, "drops": "kept", "psi_max": 0.117, "null_margin": 0.1226, "top_corr": "util_revolving 0.9998"},
{"name": "noise_a", "type": "numeric", "missing": 0.0002, "iv": 0.0155, "drops": "iv+noise", "low_iv_orgs": 3, "psi_max": 0.0943, "null_margin": 0, "sentinel_hits": 4},
{"name": "fraud_score_v2", "type": "numeric", "missing": 0, "iv": 0.0147, "drops": "iv+noise", "low_iv_orgs": 3, "psi_max": 0.3709, "psi_unstable_orgs": 3, "null_margin": -0.0018},
{"name": "bal_change_3m", "type": "numeric", "missing": 0.0691, "iv": 0.0126, "drops": "iv+noise", "low_iv_orgs": 3, "psi_max": 0.1035, "null_margin": -0.0046, "sentinel_hits": 1339},
{"name": "const_flag", "type": "categorical", "missing": 0, "drops": "constant"}
… 12 more features: dti, inq_6m, income_monthly, employment_months, social_score, app_channel, residence_type, age, loan_amount, noise_b, device_os, term_months
],
"survivors": ["bureau_score", "dti", "inq_6m", "util_revolving", "util_revolving_pct", "overdue_days_m3"],
"survivor_count": 6,
"correlated_pairs": [
{"a": "util_revolving", "b": "util_revolving_pct", "r": 0.9998}
],
"noise_screen_ran": true,
"flags": [
{"id": "F1", "severity": "medium", "category": "psi", "message": "The PSI rule needs 6 unstable organizations to drop a feature, but the modeling sample has only 3 - the PSI step cannot drop anything as set.", "features": []},
{"id": "F2", "severity": "high", "category": "leakage", "message": "2 feature(s) have overall IV above 0.5 - bureau_score (0.5756), overdue_days_m3 (16.2826). IV that strong usually means the feature was observed after the outcome or encodes the label.", "features": ["bureau_score", "overdue_days_m3"]},
{"id": "F3", "severity": "medium", "category": "missing", "message": "Sentinel values were treated as missing in feature(s) that also hold other negative numbers, so a real value may have been blanked: noise_a, bal_change_3m.", "features": ["noise_a", "bal_change_3m"]},
{"id": "F4", "severity": "low", "category": "sample", "message": "1 constant column(s) dropped before screening: const_flag.", "features": ["const_flag"]},
{"id": "F5", "severity": "medium", "category": "psi", "message": "Unstable month over month in every modeling organization, yet kept by the PSI rule as set: fraud_score_v2 (max PSI 0.3709).", "features": ["fraud_score_v2"]},
{"id": "F6", "severity": "medium", "category": "corr", "message": "1 pair(s) above the correlation threshold were left in place because both sides are among the top 20 features by IV: util_revolving ~ util_revolving_pct (r 0.9998).", "features": ["util_revolving", "util_revolving_pct"]}
],
"browser_verdict": "revise",
"method_notes": [
"IV bins: best-first decision tree, at most 5 leaves, 1% minimum leaf, missing as its own bin (toad dt recipe, re-implemented).",
"Noise screen: univariate null-IV test against label permutations - a stand-in for the reference's LightGBM null-importance gain, which does not run in a browser.",
"Correlation drop keeps the higher-IV feature of each pair (the reference uses LightGBM gain).",
"PSI: month-over-month within each organization, 10 quantile bins of the earlier month plus a missing bin; values above the earlier month's maximum go to the top bin."
]
}
The reply (data.output.output, parsed), as the model returned it. The lines starting
with … mark trimmed array entries; every value shown is real:
{
"task": "review",
"verdict": "revise",
"headline": "This screen needs one leakage check before modeling: overdue_days_m3's near-perfect IV must be confirmed as application-time before the survivors go to binning.",
"tldr": [
"Answer: dti, inq_6m and bureau_score are safe for the scorecard today; hold overdue_days_m3 until its timing is confirmed, and use only one of util_revolving/util_revolving_pct.",
"6 survivors from 19,374 modeling rows across 3 organizations (2,122 bads, 10.95% bad rate), with partner_d's 2,899 rows correctly held out of sample.",
"overdue_days_m3 (IV 16.28, null_margin 16.27) is too strong to trust without a timing check; bureau_score's IV 0.58 is plausible for a bureau score and is cleared.",
"PSI and correlation both have configuration gaps: psi_max_orgs=6 can never fire with only 3 organizations, and top_n_keep=20 equals the full feature count so the correlation guard never engages.",
"Sentinel handling likely masks real values in bal_change_3m and noise_a - all of their recorded missingness is sentinel hits, though both are already IV-dropped so it costs nothing here."
],
"step_reviews": [
{"step": "sample",
"verdict": "agree",
"note": "22,273 rows in, no invalid labels/dates, no duplicates. 2,899 partner_d rows held out. Modeling sample is 19,374 rows, 2,122 bads (10.95%) across bank_a/bank_b/broker_c - plenty of bads for a stable screen."},
{"step": "months",
"verdict": "agree",
"note": "12 months, none dropped. Monthly rows run 1,575-1,661 and bads 157-191, comfortably above min_ym_bad=10 and min_ym_total=500 - the calendar is sound."},
{"step": "missing",
"verdict": "adjust",
"note": "The 0.6 cutoff correctly caught only social_score (0.7588 missing). But bal_change_3m's whole missing share (0.0691) and noise_a's (0.0002) equal their sentinel_hits exactly - every 'missing' there is a sentinel replacement, not a true gap; review the sentinel list."},
{"step": "iv",
"verdict": "agree",
"note": "overall_iv=0.1 with org_iv=0.1/max_low_iv_orgs=2 correctly dropped 13 weak features and kept the 6 survivors plus bureau_score and overdue_days_m3, which the leakage check (F2) handles separately."},
{"step": "psi",
"verdict": "question",
"note": "psi_max_orgs=6 exceeds the 3 modeling organizations that exist, so the rule structurally cannot drop a feature (F1); step_counts.psi=0 confirms it never has. fraud_score_v2 is unstable in all 3 orgs (PSI 0.3709) yet survived (F5)."},
{"step": "noise",
"verdict": "agree",
"note": "The null-IV permutation screen correctly caught both deliberately noisy probes, noise_a and noise_b (null_margin ~0 and -0.0126), a good sanity check that the screen works as intended."},
{"step": "corr",
"verdict": "question",
"note": "top_n_keep=20 equals features_total=20, so the 'both in the top N' protection is trivially true for every pair; util_revolving/util_revolving_pct (r=0.9998) were kept together with no real prioritization (F6)."}
],
"feature_calls": [
{"feature": "overdue_days_m3",
"call": "investigate",
"engine": "kept",
"iv": 16.2826,
"reason": "IV 16.2826 with null_margin 16.2709 is far beyond what an application-time feature should show; the name suggests a days-overdue window that may overlap the 12-month performance period used to build the target. Hold for a timing check before scorecard use.",
"ref": "F2"},
{"feature": "bureau_score",
"call": "keep",
"engine": "kept",
"iv": 0.5756,
"reason": "IV 0.5756 (null_margin 0.5633) is high but a genuine bureau score can legitimately reach that on its own; no leakage indicator beyond IV strength. Safe for the scorecard.",
"ref": "F2"},
{"feature": "util_revolving",
"call": "keep",
"engine": "kept",
"iv": 0.14,
"reason": "IV 0.14, the stronger side of a near-duplicate pair (r 0.9998 with util_revolving_pct). Keep this one.",
"ref": "F6"},
{"feature": "util_revolving_pct",
"call": "drop",
"engine": "kept",
"iv": 0.1393,
"reason": "Overriding the corr step: it kept this feature only because top_n_keep=20 equals the full feature count, so the correlation guard never actually screens anything here. r=0.9998 with util_revolving means it adds no information; drop the redundant copy.",
"ref": "F6"}
… 7 more calls: dti keep, inq_6m keep, noise_a drop, bal_change_3m drop, const_flag drop, fraud_score_v2 drop, income_monthly investigate
],
"leakage_suspects": [
{"feature": "overdue_days_m3",
"evidence": "IV 16.2826, essentially unchanged after label permutation (null_margin 16.2709) - far outside the range a legitimate application-time feature should show.",
"check": "Confirm overdue_days_m3 is computed strictly from data timestamped before apply_date and not from the 12-month performance window; rebuild it from pre-application snapshots only and rescore its IV."}
],
"threshold_advice": [
{"parameter": "sentinels",
"current": "-1, -999, -1111",
"suggested": "-999, -1111",
"why": "bal_change_3m's missing share (0.0691) equals its sentinel_hits (1339) over modeling_rows exactly, and noise_a's missing (0.0002) equals its 4 sentinel_hits the same way - every recorded 'missing' value in both is a sentinel hit. Confirm -1 is not a legitimate small value before keeping it as a sentinel."},
{"parameter": "psi_max_orgs",
"current": 6,
"suggested": 2,
"why": "Only 3 modeling organizations exist (bank_a, bank_b, broker_c), so psi_max_orgs=6 can never be reached (F1); fraud_score_v2 is unstable (PSI 0.3709) in all 3 yet survived (F5)."},
{"parameter": "top_n_keep",
"current": 20,
"suggested": 8,
"why": "top_n_keep=20 equals features_total=20, so the correlated-pair protection always applies to any pair; util_revolving and util_revolving_pct (r=0.9998) were kept together with no real prioritization (F6)."}
],
"next_steps": [
"Confirm overdue_days_m3 is built from data timestamped strictly before apply_date, not the 12-month performance window; rebuild and rescore if it isn't.",
"Drop util_revolving_pct and keep util_revolving (r=0.9998, marginally higher IV) before binning.",
"Lower psi_max_orgs to 2 of 3 organizations and rerun PSI so instability like fraud_score_v2's can actually drop a feature.",
"Lower top_n_keep well below features_total (e.g. 8) so the correlation guard stops protecting every pair by default.",
"Review the sentinel list (-1, -999, -1111) against bal_change_3m and noise_a specifically before trusting their missing shares.",
"Take bureau_score, dti, inq_6m and util_revolving into scorecard binning now; hold overdue_days_m3 pending the leakage check.",
"Revisit income_monthly after cleaning its sentinel-driven missingness (943 hits) - its IV of 0.0548 is borderline and worth a second look."
],
"memo": [
"The retail-instalment-2025 sample has 22,273 loan-level rows, no invalid labels or dates and no duplicates. 2,899 rows from partner_d are held out of sample entirely; the modeling sample is bank_a, bank_b and broker_c, 19,374 rows, 2,122 bads (10.95% bad rate) over 12 months, none dropped. Monthly bads run 157-191 and rows 1,575-1,661, comfortably above the min_ym_bad and min_ym_total floors.",
"Of 20 features, const_flag was removed as constant, social_score failed missing-rate (75.88%), 13 failed IV, 8 failed noise (incl. probes noise_a/noise_b), 0 failed PSI/corr - not for lack of instability or correlation, but because psi_max_orgs=6 exceeds the 3 organizations and top_n_keep=20 equals features_total. fraud_score_v2 and util_revolving_pct both slipped through that way."
… 2 more paragraphs: leakage and stability concerns, survivors and next steps
],
"prescan_responses": [
{"ref": "F1",
"verdict": "confirmed",
"note": "Only 3 modeling organizations exist against psi_max_orgs=6 - the rule needs more unstable organizations than exist, so PSI can never drop a feature as configured (step_counts.psi=0 confirms it never has)."},
{"ref": "F2",
"verdict": "confirmed",
"note": "overdue_days_m3 (IV 16.2826, null_margin 16.2709) is implausibly strong for an application-time feature and is treated as a leakage suspect. bureau_score (IV 0.5756, null_margin 0.5633) is a standard bureau score that can legitimately reach this IV on its own - cleared, kept as a survivor."},
{"ref": "F3",
"verdict": "confirmed",
"note": "bal_change_3m's missing share (0.0691) exactly equals its sentinel_hits over modeling_rows (1339/19374), so every 'missing' value there is a sentinel replacement, not a true gap; noise_a is the same (4/19374). Both are already IV-dropped so there is no scorecard impact, but the sentinel list should be reviewed."}
… F4 confirmed, F5 confirmed, F6 confirmed
]
}
Note what the review did that the screen could not: it kept bureau_score (a bureau
score can reach IV 0.5 on its own) while holding overdue_days_m3 as a leakage suspect,
called util_revolving_pct drop although the screen kept it, and flagged the
two rules that could not fire as set. It changed nothing: the screen's result is still what
engine says.
Truncation and partial results
When the balance sits between min_credits and hold_credits, the run is not
refused: it executes with a reduced output cap and reports "truncated": true, in the
done event of /run-stream and on the job from GET /jobs/{id}.
What you hold is then a prefix of the reply, and it will not parse as it stands. The web page closes
the cut-off JSON (Recon.closeJson in recon.js: close an open string, drop
a dangling comma or key, close every open array and object, and if that still does not parse, cut
back to the previous comma and try again), parses what is left, and shows the sections that arrived
as "N of 11 sections recovered". A stream ended early error gets the same
treatment on the deltas received so far.
The keys arrive in contract order, so a cut costs the tail first: memo and
prescan_responses go first, then next_steps and
threshold_advice. A truncated review can therefore carry a verdict while
the flag answers that justify it are missing, and it will fail the flag and verdict-floor checks
above. Never file one as final: check the flag, top up, and resubmit with the attempt suffix on the
Idempotency-Key incremented (iv-desk:review:<hash>:a2).
If a complete reply will not parse as one JSON object, the page retries once, as the next attempt,
with a retry_note saying what was wrong and asking for only the JSON object for task
review, no prose, no code fences, every key present. Do the same: keep the hash of the
original body, bump the attempt, add retry_note.