ValueFun.ai

In the beginning was the Value.

Value functions compress the future. They reduce variance, guide policy, and turn noisy experience into direction.

Welcome to ValueFun — a small meditation on value, judgment, reinforcement learning, and the long horizon of intelligence.

May you and your model enjoy high explained variance in value prediction — more so than low perplexity in policy.
I. Policy and Prophecy

The Gospel of Value

In reinforcement learning, the policy wanders through the wilderness of action, reward, uncertainty, and consequence.

The value function is its prophecy.

It sees, however imperfectly, the return that lies beyond the next token, the next action, the next step through the desert.

A good value function gives the policy a shortcut through uncertainty. It reduces the variance of our gradients. It turns experience into direction. It transforms wandering into search.

We believe value prediction is not merely a training trick. It may be one of the missing keys to more advanced intelligence.

For what is intelligence, if not the ability to perceive what matters before the reward is fully revealed?

II. The Long Horizon

The policy samples. The value judges.

The policy says:
“Which action shall I take?”

The value function says:
“Behold the long horizon.”

The policy explores the many paths. The value whispers which paths may lead to reward.

And when the gradient is noisy, the value function becomes a lamp unto the optimizer’s feet, and a light unto its trajectory.

III. Beyond Imitation

Perplexity is not wisdom.

Low perplexity is useful. But perplexity is not wisdom.

A model may predict the next token and still fail to understand the consequence of its continuation.

A model may imitate the distribution and still lack judgment.

We seek models that do more than speak fluently. We seek models that evaluate, anticipate, correct, and care about outcomes.

The future of intelligence may not belong only to better policies, larger datasets, lower losses, or lower perplexities.

It may belong to systems with better value.

IV. The Blessing

May your rollouts be stable.

And above all:

May your value predictions explain the world before your policy acts within it.