Skip to content

Pull requests: EvolvingLMMs-Lab/lmms-eval

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[Model] Add VQToken integration for LLaVA-OneVision
#1430 opened Aug 11, 2026 by Hai-chao-Zhang Loading…
3 of 7 tasks
feat(tasks): add C4 Bench
#1429 opened Aug 10, 2026 by sci-m-wang Loading…
2 of 7 tasks
fix(vllm): preserve per-request sampling parameters
#1428 opened Aug 9, 2026 by turturturturtur Contributor Loading…
feat(tasks): add owner-aligned MicroVQA benchmark
#1427 opened Aug 9, 2026 by turturturturtur Contributor Loading…
fix(openai): preserve per-request sampling parameters
#1426 opened Aug 9, 2026 by turturturturtur Contributor Loading…
[Model] Add Mistral3-VL simple model
#1425 opened Aug 7, 2026 by thanksmumu Loading…
1 task done
feat(agentic): generalize the environment interface beyond ViZDoom
#1419 opened Jul 31, 2026 by Luodian Contributor 2/2 Draft
3 of 7 tasks
feat: add agentic game-loop evaluation (generate_until_game)
#1418 opened Jul 31, 2026 by Luodian Contributor 1/2 Loading…
4 of 7 tasks
ViZDoom
#1368 opened Jun 21, 2026 by pufanyi Collaborator Loading…
Add Qwen-native JSON coordinate variants for pointing tasks
#1361 opened Jun 5, 2026 by njb-nvidia Contributor Loading…
Feat/ollama model
#1322 opened May 5, 2026 by eliasubz Loading…
1 of 7 tasks
feat: vLLM-Omni for video generation models
#1314 opened Apr 27, 2026 by pufanyi Collaborator Draft
fix(evaluator): auto-init gloo process group for multi-rank launches
#1306 opened Apr 23, 2026 by Luodian Contributor Draft
2 tasks
fix(api/task): guard empty results list in process_results
#1305 opened Apr 23, 2026 by Luodian Contributor Draft
2 tasks
Fix missing Task import for type annotation in evaluator
#1291 opened Apr 10, 2026 by luv-oct22 Loading…
2 tasks
feat: add VBench video generation evaluation benchmark
#1271 opened Mar 26, 2026 by Luodian Contributor Loading…
3 tasks
feat: add MiniMax as LLM judge provider (default model: MiniMax-M3)
#1263 opened Mar 22, 2026 by octo-patch Loading…
3 tasks done
ProTip! Filter pull requests by the default branch with base:main.