Replies: 1 comment
|
Thanks @Astro-Han for putting this forward — it's a valuable and worthwhile effort. Publishing a reproducible Maka run against the FrontierHarness suite, with pinned config and disclosed environment differences, is exactly the kind of grounded, community-verifiable evidence the project benefits from. I also appreciate that the scope is kept deliberately small and the budget framed as a planning cap rather than a promise. I'd like to help. Happy to take on (including but not limited to):
Let me know which piece would be most useful to pick up first. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I'd like to organize a community run of Maka on Runta's FrontierHarness Eval and publish the results alongside the existing harness baselines.
The scope is deliberately small: run Maka only, on the published 30-task suite. We would use the published results for other harnesses as reference points, rather than pay to rerun the whole comparison.
Proposed setup
Budget
I propose a $100–150 total spending cap for the first Maka run, including setup and runtime costs. This is a planning limit, not a measured Maka cost or a guarantee that the run will fit.
For scale, summing the published standardized task costs gives about $69 for Codex and $44 for Pi, each with all 30 task costs present. Those figures include failures, use the benchmark's pricing/cache normalization, and exclude runtime charges; they are not live provider invoices. Published data.
Runta currently offers $50 in trial runtime credit for new accounts, which may cover the execution environment if available to the account we use. Fireworks API usage is separate. Runta pricing.
What we would publish
Maka's pass rate, task-level outcomes, token costs, runtime, and failure analysis, together with pinned configuration and reproducible instructions. We should retain formal failures and disclose missing evidence or infrastructure-invalid trials.
This would be a later, independent Maka run, not a simultaneous reproduction of the original experiment. The comparison should clearly disclose version, serving, and environment differences, especially for latency and cache-sensitive cost measurements. Both the actual spend and the benchmark-standardized cost would be useful.
Join in
Would anyone like to help with:
If you're interested, please reply with the part you'd like to help with. We can agree on an operator, available credits, and the expense arrangement here before starting paid runs. Please don't post API keys.
Any personal cost-sharing would need a clearly identified organizer and an expense record; this thread is a proposal, not an ASF or Apache Maka fundraising campaign. If support is arranged as official project sponsorship, we should coordinate the appropriate route with the mentors.
All reactions