Reading ListA running collection of original blogs and articles I found worth keeping. Each card links straight to the source. Learning to Replicate Expert Judgment in Financial Tasks thinkingmachines.ai Putting Task Expertise into RL Achieves State-of-the-Art Performance on Text-to-SQL thinkingmachines.ai 从OPD与反向KL的关系到OPD的两种形态以及路线之争 zhuanlan.zhihu.com 量化:用更少的显存跑更大的模型 inferloop.dev Harness Engineering for Self-Improvement lilianweng.github.io LoRA Without Regret thinkingmachines.ai On-Policy Distillation thinkingmachines.ai Interaction Models: A Scalable Approach to Human-AI Collaboration thinkingmachines.ai Reward hacking is swamping model intelligence gains cursor.com