On September 30, webAI launched the TwIL-LM3-Pro, a language model with 3.66 billion parameters specialized in formal reasoning. In tests published by the developer, the system achieved performance equivalent to Qwen3 8B on its main logic metric, despite having less than half the parameters.
On the aggregate formal logic metric used by webAI, TwIL-LM3-Pro scored 0.5539, versus 0.5336 for Qwen3 8B. The difference, however, fell within the tests' sampling noise, and the company itself states that the result should be interpreted as parity, not as a victory over the larger model.
The advance is clearer in comparison with IBM Granite 4.2 3B, the model that served as the basis for the development. The score went from 0.4313 to 0.5539, a relative improvement of about 28%. TwIL-LM3-Pro was also ahead of the public version of VibeThinker-3B in the six formal logic tasks compared by the company.
On broader benchmarks, the model achieved 95.4% on the logic portion of BIG-Bench Hard and 95% on SVAMP, aimed at written math problems. On GSM8K, it recorded 94.3%, close to the 95.7% obtained by Qwen3 8B.
The results, however, do not indicate that the smaller model has reached Qwen3 8B in general capability. On average across ten benchmarks outside its specialized domain, TwIL-LM3-Pro scored 0.7901, while Qwen3 8B reached 0.8493. The competitive advantage appears mainly in the logic tasks for which the model was trained.
2.09 GiB version can run locally
TwIL-LM3-Pro was developed with a post-training sequence that includes supervised fine-tuning, checkpoint combination, weight interpolation, and reinforcement learning. The specialization covers tasks such as logical inference, translation from natural language to formal logic, and proof analysis.
In addition to the full weights in BF16, which occupy 6.82 GiB, webAI made quantized versions available. The file Q4_K_M occupies 2.09 GiB and, according to the documentation, can be run on CPU or with 4 GB of VRAM, reducing the need for cloud infrastructure for certain uses.
There is an important caveat: the released benchmarks were run with the weights in BF16. webAI has not yet measured how much quantization affects the accuracy of the Q4 version, so the published results cannot be directly attributed to the smaller file.
The model was released under the webAI Non-Commercial License v1.0, which restricts its commercial use. The full weights and GGUF versions are publicly available for local execution and testing.



