inspiration

cloud coding assistants assume a paid API and a stable network. many African students work on ordinary 8 GB laptops. we wanted one local model that does two jobs: write and repair code, and act as an agent with real one-line tool calls. then i chose this model because the vibe-thinker as base because it was already a fairly good model on competitive programming and maths/STEM. except one issue- the model is bad at agentic coding because of over-reasoning about calling tools rather than just calling it - which is why i took it up and started fine tuning this model to run locally with good agentic capabilities

what it does

VON-Coder-3B is a model capable of coding near opus4.5 level as base was below opus but since our training and evaluation proved to improve it- we estimate it around that range

we started from WeiboAI/VibeThinker-3B. On the same EvalPlus 0.3.1 HumanEval bank, the model lifts HumanEval from 0.866 to 0.921 and HumanEval+ from 0.817 to 0.884.

on the 32-row tool probe the model emits a valid one-line tool call with short think on 32/32. the base emits 0/32. during GRPO we trained it to prefer shorter and correct tokens- where base on our test fit into 32k ctx 0/32 times, and our model 32/32 times. profiler numbers from our run: 5.29 tok/s, peak RSS 3312.59 MB, Sperf 35.27, Seff 52.68.

how we built it

base: WeiboAI/VibeThinker-3B. we designed group-conditioned adaptive LoPD: keep GRPO when a group still has a verified correct rollout; distill with LoPD only on failed rollouts (50/50 bidirectional KL) with no EMA teacher. then LoRA and weight modification, merge, GGUF Q8_0., trained for tool calls, shorter reasoning between correct rollouts in groups got more reward, therefore reduced over-reasoning, 210 Groups for rl and distillation, SFT on ~108k agentic/coding rows, with proper language distribution for better generalization, we went from SFT -> RL + grpo(using the distillation method already specified), -> lora and quantization after, then evaluation on 80hours+ rtx pro 6000 and 40hrs+ of rtx a6000 on vast cloud

challenges

fitting a useful 3B agentic coder on the 8 GB CPU profile under a compute constraint, keeping thinking short enough for tools, and proving the coding lift on the same harness as the base. also encountered lots of issues when trying to reduce vibe thinker over-reasoning on hard coding tasks, i increased the rewards for shorter reasoning at almost every 20 group rollouts😅- this solved it then due to the previous challenge of over-reasoning, this model struggled to call tools which would have been bad for agentic tasks, i had to upweight tokens for tool calls in sft, mix more agentic tasks during rollouts and even during lora, modify the weights by performing lora on a sweet spot where the model apparently had coverage for tool calling although wasnt in topk, so i noticed it had the tendency when i gave more coverage for agentic tasks and i did many values trying to find the spot until it finally clicked and i proceeded with that model which is now the submitted one.

what i learned

it was due to this project that i was able to do a research on the distillation method that would work best for us and thats why i started: group-conditioned adaptive LoPD: Standard GRPO needs at least one verified-correct rollout in the group. An all-failed group has no trustworthy positive contrast. Blind distillation on every token is also wrong, because a privileged teacher can overwrite a group that already has a real win. So we keep GRPO on whenever the group still has a verified correct rollout, and we apply LoPD only to failed rollouts that need a privileged teacher. For group g: L_student = c_GRPO(g) * L_GRPO + alpha * lambda_g * L_LOPD(eligible(g)). c_GRPO is 0 only if every rollout failed, else 1; it is not (1 - lambda_g). Eligibility: 0 correct -> LoPD on all failures; 1 correct -> GRPO on the full group and LoPD only on the failures; 2+ correct -> GRPO only. lambda_g rises when reward evidence is weak and teacher-verifier agreement is strong. the teacher is not an EMA of the student. It is a frozen policy that rescores the student's actual prefixes with teacher-only latent context that is not available at laptop inference. Deadline mix is 0.5 * KL(teacher || student) + 0.5 * KL(student || teacher) (top-64 logits plus a tail bucket). LoPD (arXiv:2608.13040) supplies the latent-context substrate; I-SDPO (arXiv:2608.12957) supplies shared group routing. we keep GRPO at coefficient 1, replace the EMA teacher, fix 50/50 bidirectional KL, and add the success-count gate. Correctness comes first; during GRPO we apply length penalties only after a response is verified correct or structurally valid, so the model is trained to prefer shorter and correct tokens.

Built With

  • agentic
  • coding-assistant
  • gguf
  • hugging-face
  • llama.cpp
  • offline-ai
  • peft
  • python
  • pytorch
  • transformers
  • vast
Share this project:

Updates

Submission history