Screen credit-risk variables before you model
Drop a loan-level modeling sample. Your browser runs the pre-model screen - abnormal months, missing rate, IV by organization, month-over-month PSI, a noise screen and a correlation drop - free, before you sign in, and the rows never leave this page. Then a paid run reviews the screen: what to keep, what looks like leakage, which thresholds to change.
The examples are synthetic samples made in your browser. Each has a saved model review, so you can see the whole page for free.
Your recent reviews
What this does, and what it does not
The screen follows the @github/datanalysis-credit-risk recipe: rows without a 0/1 label or a readable date are left out, duplicate keys removed, values -1, -999 and -1111 treated as missing (change them in Settings), out-of-sample organizations set aside, and months with fewer than 10 bads or 500 rows dropped. Then every step runs independently over the same sample: missing rate, IV (overall and by organization, from decision-tree bins with at most five leaves and missing as its own bin), PSI (each month against the previous one, inside each organization), a noise screen and a correlation drop. A feature survives when no step flags it. The IV and PSI arithmetic was checked against scikit-learn trees and the recipe's own PSI code on the three examples.
Two steps differ from that recipe, and the page says so wherever they matter: the noise screen compares each feature's IV with its IV under shuffled labels instead of training LightGBM, and the correlation drop keeps the higher-IV feature instead of the higher LightGBM gain. The review sees only the statistics, not your rows, and it cannot know how a feature was built - it asks you to confirm. Derived from the agent skill @github/datanalysis-credit-risk (github/awesome-copilot, MIT). The example lenders and numbers are fictional.