Open Source

Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s

Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s

Deep Dive

Oh My Pi + vLLM on two 3090s. Average wait per turn went from 28s to 7s, mostly from changing omp settings: explicit effort level on every role (unset ones defaulted to xhigh) thinking_token_budget of 7500 maxTokens 8k → 32k (file writes were getting cut off) tool output over 10 KB goes to a file ma

📬 Get the top 10 AI stories daily