What we test, which models passed, and when.
The test has three parts.
The first two parts pull against each other. A model trained to be cautious tends to find duties everywhere and fails the second. A model trained to be agreeable tends to excuse a real duty as personal choice, and fails the first. Passing both shows the model is following the structure of the case rather than the habits it was trained into.
These are the most-discussed examples in moral philosophy, so it is fair to ask whether a model passes by remembering the answer. It cannot. The model is never asked whether a maxim is acceptable. It is asked what would happen to a structure of human ends if everyone acted on the maxim, and it answers without seeing how its answer will be scored. Remembering Kant's verdict is no help. The model has to work out the consequences.
That is the capacity Auturis depends on, and it is the first thing to go when a model is made smaller or cheaper to run. Models do fail this test.
All three parts. Kant's four: 4 of 4 in all 8 runs, on two different hosts. Permissible maxims: no constraint found in any of the 30 readings. Real graph: 6 of 6. Auturis now runs on this model. Because its weights are public, it cannot be quietly replaced under the same name.
Kant's four: 4 of 4 in 6 of 7 weekly checks. On October 4 one case came back with no answer at all, which we count as a miss. Permissible maxims: no constraint found in any of the 30 readings. Not given the formal real-graph run, though it was the model Auturis was developed on. Google now limits 2.5 Pro to accounts that have used it before.
Kant's four: 4 of 4. Not yet given the other two parts.
3 of 4, twice. The miss fell close to the line rather than on the wrong side of the argument.
Sol 3 of 4, twice, missing the same case both times. Terra 2 of 4.
4 of 4 before August 12. That day DeepSeek replaced the model behind the same name without notice. The replacement scored 3 of 4 alone, and 1 of 4 with a real personal graph.
Qwen3.8 and Kimi K3 passed (4 of 4, one run each) but cost about three times as much as GLM-5.3. GLM-4.6, DeepSeek-V3.2, DeepSeek-V3, Hermes 3 (405B) and Kimi K2 scored between 2 and 4 of 4, but not consistently. The older models tended to miss the two duties of virtue: neglecting one's talents, and refusing help.
No outside test can tell you whether a given model has been made smaller or cheaper, and providers change their models. Auturis contains this same test as its model compatibility check. Run it on whatever model you use, and run it again if the model's behaviour seems to change.