Guides

Read options

Four optional fields change how OpenJev reads the answers. Leave them out and a request behaves exactly like Jev.

Overview

FieldValuesWhat it doesCost
steps1–8, default 1Denoise steps per read. More steps let the answers settle against each other.Same tokens, more GPU time
samples1–32Read N times with different noise and average the distributions. Replaces the automatic re-reads.N × input tokens
think0–4096Let the model write a thought of up to this many tokens, then read the answers after it.Input tokens twice, plus the thought
sequentialtrueAnswer long question lists in chunks, in order, each chunk seeing the answers already chosen.One read per chunk, one after another

These fields are OpenJev additions and come from the example server in vLLM PR #57250. The TypeSafe SDKs never send them, so SDK code is unaffected.

steps

A read normally takes one denoise step, so every answer is read at once. With steps above 1 the model denoises the answer slots several times, and each answer can settle against the others. Try it when questions are related, for example a category followed by a subcategory. The token cost is the same; each extra step adds GPU time.

samples

samples: 8 reads the same request eight times with different noise and averages the eight distributions before picking choice and score and computing confidence. It gives smoother probabilities, which helps when you set thresholds. Each read is billed.

think

think lets the model reason in text before it answers. OpenJev generates a thought of at most think tokens, then reads the answers with the thought in context. It helps with arithmetic and multi-step problems:

Request
{
  "model": "openjev-latest",
  "state": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball.",
  "think": 512,
  "questions": {"ball_5c": {"type": "noul", "instructions": "The ball costs 5 cents."}}
}

The budget is a hard cap: a thought that reaches it is cut off mid-sentence. Give multi-step problems 512 or more. The input is read twice (once to think, once to answer), and the thought tokens are reported as usage.output_tokens. Both count toward your System One quota.

sequential

Large question sets are answered in parallel chunks, and a chunk can't see the others' answers. sequential: true reads the chunks in order instead, and each chunk sees the answers already chosen. Use it when later questions depend on earlier ones, for example tagging every word of a sentence. It is slower, because the chunks can't run in parallel.

Images

steps and samples work with images. think and sequential need a text-only state; combining them with images returns 422.

Try it

Every option is in the Advanced panel of the Raw playground, with a preset for each.