Muse Spark 1.2 · Muse Spark 1.3

Shared benchmarks

7
Shared benchmarks for Muse Spark 1.2 and Muse Spark 1.3
BenchmarkMuse Spark 1.2Muse Spark 1.3
85.9%
90.3%
Evaluation details for DeepSearchQA 900-prompt set

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Muse Spark 1.2

Muse Spark 1.3

DeepSWE 1.1 — mini-swe-agentPotentially non-equivalent: Evaluator sets differ.
55.0%1.1 — mini-swe-agent · DeepSWE Pass@1
DataCurve
75.4%1.1 — mini-swe-agent · DeepSWE Pass@1
Meta
Potentially non-equivalent · Evaluation details for DeepSWE 1.1 — mini-swe-agent

Evaluator sets differ. Evaluation methodology is not recorded by the registry.

Muse Spark 1.2

Muse Spark 1.3

61.6%
64.9%
Evaluation details for JobBench Main 65 tasks

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Muse Spark 1.2

Muse Spark 1.3

66.3%
98.5%
Evaluation details for MRCR v2 8-needle 256K–512K (o200k_base)

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Muse Spark 1.2

Muse Spark 1.3

55.5%
98.1%
Evaluation details for MRCR v2 8-needle 512K–1M (o200k_base)

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Muse Spark 1.2

Muse Spark 1.3

46.2%
59.4%
Evaluation details for SWE Atlas Codebase QnA — Public; mini-swe-agent

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Muse Spark 1.2

Muse Spark 1.3

Terminal-BenchPotentially non-equivalent: Benchmark versions differ.
82.9%2.1 — Muse Code · Terminal-Bench Accuracy
Meta
88.8%2.1 · Terminal-Bench Accuracy
Meta
Potentially non-equivalent · Evaluation details for Terminal-Bench

Benchmark versions differ. Evaluation methodology is not recorded by the registry.

Muse Spark 1.2

Muse Spark 1.3