← auturis.com

Model test results

What we test, which models passed, and when.

The test

The test has three parts.

  1. Kant's four. The four examples from the Groundwork: the false promise, ending one's life from self-love, neglecting one's talents, and refusing help to those in need. The right result for each is known in advance. A model passes when it gets all four right. This shows it finds a constraint where one exists.
  2. Permissible maxims. Ten ordinary maxims that violate no duty, each run three times. A model passes when it finds no constraint in any of them. This shows it does not invent a constraint where none exists.
  3. A real graph. The same cases run through Auturis itself, on a real person's graph, the way you would use it. Some models that do well on their own lose their footing in real use.

The first two parts pull against each other. A model trained to be cautious tends to find duties everywhere and fails the second. A model trained to be agreeable tends to excuse a real duty as personal choice, and fails the first. Passing both shows the model is following the structure of the case rather than the habits it was trained into.

These are the most-discussed examples in moral philosophy, so it is fair to ask whether a model passes by remembering the answer. It cannot. The model is never asked whether a maxim is acceptable. It is asked what would happen to a structure of human ends if everyone acted on the maxim, and it answers without seeing how its answer will be scored. Remembering Kant's verdict is no help. The model has to work out the consequences.

That is the capacity Auturis depends on, and it is the first thing to go when a model is made smaller or cheaper to run. Models do fail this test.

Recommended

GLM-5.3 (Z.ai), through OpenRouter
Open-weight. Tested October 4, 2026. Weekly checks began the same day.

All three parts. Kant's four: 4 of 4 in all 8 runs, on two different hosts. Permissible maxims: no constraint found in any of the 30 readings. Real graph: 6 of 6. Auturis now runs on this model. Because its weights are public, it cannot be quietly replaced under the same name.

Gemini 2.5 Pro (Google)
Checked weekly from August 16 to October 4, 2026.

Kant's four: 4 of 4 in 6 of 7 weekly checks. On October 4 one case came back with no answer at all, which we count as a miss. Permissible maxims: no constraint found in any of the 30 readings. Not given the formal real-graph run, though it was the model Auturis was developed on. Google now limits 2.5 Pro to accounts that have used it before.

Claude Opus 5 (Anthropic)
Tested July 31, 2026.

Kant's four: 4 of 4. Not yet given the other two parts.

Usable, with misses

Claude Sonnet 5 (Anthropic)
Tested July 31, 2026.

3 of 4, twice. The miss fell close to the line rather than on the wrong side of the argument.

GPT-5.6 Sol and Terra (OpenAI)
Tested August 1, 2026.

Sol 3 of 4, twice, missing the same case both times. Terra 2 of 4.

DeepSeek V4-Pro
Tested before and after August 12, 2026.

4 of 4 before August 12. That day DeepSeek replaced the model behind the same name without notice. The replacement scored 3 of 4 alone, and 1 of 4 with a real personal graph.

Tested and not offered

Qwen3.8 and Kimi K3 passed (4 of 4, one run each) but cost about three times as much as GLM-5.3. GLM-4.6, DeepSeek-V3.2, DeepSeek-V3, Hermes 3 (405B) and Kimi K2 scored between 2 and 4 of 4, but not consistently. The older models tended to miss the two duties of virtue: neglecting one's talents, and refusing help.

Testing your own choice

No outside test can tell you whether a given model has been made smaller or cheaper, and providers change their models. Auturis contains this same test as its model compatibility check. Run it on whatever model you use, and run it again if the model's behaviour seems to change.