跳转至

比较有限采样下的 VQE 目标估计

优化过程中出现一个低能量,可能有一部分来自采样波动。本项目比较四种目标估计策略,独立复测候选参数,再用小规模模拟器精确检查所选参数的能量。你需要区分估计误差与参数质量,并算清得到这些结果用了多少次执行。

先完成 VQE 与能量估计、基线分析和实验设计。按安装说明准备环境,在仓库根目录运行。默认使用一个量子比特、四种策略、三轮配对重复,每种策略只做一步 SPSA。实验在本地 CPU 上采用理想状态演化,测量的不确定性只来自有限次采样。

推导一个可以核对的参照值

使用 H = 0.1 I + 0.7 X − 0.4 Z,能量单位任意。由于 (0.7 X − 0.4 Z)² = 0.65 I,最低特征值为 E0 = 0.1 − sqrt(0.65) ≈ −0.70623。这个基态能量可以独立计算。

默认的单层线路对 |0⟩ 依次施加 RY(θ)、RZ(φ),能量为:

E(θ, φ) = 0.1 + 0.7 sin(θ) cos(φ) − 0.4 cos(θ).

把保存的 ry_0_0、rz_0_0 代入即可核对。精确求值消除了目标估计的采样误差,但不能保证优化器只走一步就找到基态。

X、Z 两项需要不同的测量基。X 测量在计算基测量前施加 H 门,Z 组无需旋转测量基。每组用计数估计取值为 ±1 的均值,采样能量为 0.1 + 0.7 mean(X) − 0.4 mean(Z)。

明确预算究竟限制什么

同一轮中,四种策略使用相同的初始参数和 SPSA 种子;不同轮使用不同的派生种子。令扰动符号 Δi ∈ {−1,+1},SPSA 估计 gi = [E(θ+cΔ) − E(θ−cΔ)]/(2cΔi),再更新 θi ← θi − a gi。本例取 a = 0.2、c = 0.15,只更新一次。初始点、正扰动点、负扰动点、最终中心点,共产生四个逻辑目标估计。

策略 一个逻辑目标估计怎样得到
exact 从模拟状态精确求值一次
sampled_single 完整测量一次,包含 X、Z 两组
sampled_fixed 完整测量两次,合并统计
sampled_adaptive 至少完整测量两次;合并标准误差达到 ≤ 0.08 时停止,否则最多测三次,同时受预算限制

每次运行最多执行 12 次目标求值,重复测量也计入其中,不要求用满。精确策略和单次采样通常各用四次,固定重复用八次,自适应重复可用八至十二次。这里相等的是求值上限,实际采样次数和总成本并不相等。

设置 shots_per_group=128、allocation="coefficient_l1" 后,一次完整采样求值的总预算为 2 × 128 = 256 次。按系数绝对值分配后,X 组得到 163 次,Z 组得到 93 次。在这种分配方式下,不能根据参数名判断每组一定获得 128 次。

优化完成后,采样策略从历史记录中选出最多两组不同的候选参数,每组独立复测两次。默认运行均有两个候选,因此每个采样策略的确认阶段使用八次 Backend 执行和 1,024 次采样。评价候选选择时,只报告搜索历史中出现过的最低采样能量,会遗漏筛选带来的偏差。

每种策略还对选中的状态做 128 次末端计算基采样。它描述 Z 基下的分布,不能单独估计 X 项。基准程序另用精确模拟器检查每个采样策略所选参数的能量;这次诊断也要计算成本,在真实硬件上不能假定有免费的精确答案。候选确认、末端采样和精确诊断都不计入优化阶段的 12 次求值上限。

运行并检查结果

python3 examples/learning/projects/finite_shot_vqe.py

打开 artifacts/finite-shot-vqe/vqe-zh.html。标准基准报告展示所选能量、配对比较、不确定性和执行成本。benchmark.json 保存 Hamiltonian、配置、所选参数、完整优化与确认记录、各测量基的计数和种子、精确检查及汇总统计。两个语言版本的报告使用同一份运行结果。

打开实验报告

下载原始数据(JSON)

{
  "artifacts": {
    "data": "artifacts/finite-shot-vqe/benchmark.json",
    "report_en": "artifacts/finite-shot-vqe/vqe-en.html",
    "report_zh": "artifacts/finite-shot-vqe/vqe-zh.html"
  },
  "ground_energy": -0.706225774829855,
  "objective_evaluation_ceiling": 12,
  "paired_repeats": 3,
  "sdk_version": "1.0.8a",
  "strategies": [
    {
      "estimator_error_rmse": null,
      "mean_exact_selected_energy": -0.12351605330345887,
      "mean_paired_exact_reference_gap": 0.0,
      "strategy": "exact"
    },
    {
      "estimator_error_rmse": 0.01341777011131602,
      "mean_exact_selected_energy": -0.1163455933580984,
      "mean_paired_exact_reference_gap": 0.007170459945360484,
      "strategy": "sampled_single"
    },
    {
      "estimator_error_rmse": 0.03697706494314943,
      "mean_exact_selected_energy": -0.11853708251671276,
      "mean_paired_exact_reference_gap": 0.004978970786746112,
      "strategy": "sampled_fixed"
    },
    {
      "estimator_error_rmse": 0.04777476599368023,
      "mean_exact_selected_energy": -0.11556891288183073,
      "mean_paired_exact_reference_gap": 0.00794714042162815,
      "strategy": "sampled_adaptive"
    }
  ],
  "total_backend_executions": 229,
  "total_shots": 26624
}

开发环境的默认运行中,十二次策略运行合计使用 229 次 Backend 执行、26,624 次采样,包含上述四类成本。精确策略所选参数的精确能量均值约为 −0.12352,单次采样策略约为 −0.11635,都明显高于基态能量。这套小配置用于学习实验流程,还没有收敛到最优解。

阅读结果时区分三个量:

  • selected_estimator_energy:采样策略独立确认后得到的能量估计,不是搜索历史里的最低采样值。
  • estimator_error:从上述估计中减去同一组所选参数的精确能量,衡量该点的估计误差。
  • paired_exact_reference_gap:从所选参数的精确能量中减去同一轮精确优化策略的所选能量,比较两次优化的参数质量。它不是与真实基态的差距。

配对差为负,只说明采样策略在这一对运行中碰巧找到了更好的参数,不能据此判断它普遍更好。默认运行中的九次采样候选选择均为 inconclusive:程序仍会返回一个候选,但确认数据没有充分区分它和竞争候选。

基准程序根据三轮结果计算样本方差、标准误差及均值的双侧 Student-t 区间。只有三个观测值时,95% 临界值约为 4.303,区间可能很宽。这些区间描述重复运行之间的变化;单个候选的测量误差回答的是另一个问题。两者都不能证明全局最优,在多个策略中挑选最有利的估计本身也会引入选择效应。

"""Compare exact and sampled VQE objectives under a paired evaluation budget."""

from __future__ import annotations

import argparse
import json
from math import sqrt
from pathlib import Path
from typing import Any, Literal

import cascaqit
from cascaqit import (
    VQE,
    HamiltonianTerm,
    OptimizerConfig,
    PauliHamiltonian,
    PauliMeasurementConfig,
    PauliX,
    PauliZ,
    SampledSelectionConfig,
    SPSAConfig,
    VQESamplingBenchmarkConfig,
)


def experiment(
    output_dir: Path, *, repeats: int = 3, shots: int = 128, budget: int = 12
) -> dict[str, Any]:
    """Reuse executed benchmark data for both language reports and the JSON file."""
    hamiltonian = PauliHamiltonian(
        hamiltonian_id="project.finite-shot.hamiltonian",
        logical_order=("q0",),
        constant=0.1,
        terms=(
            HamiltonianTerm("x", 0.7, PauliX("q0")),
            HamiltonianTerm("z", -0.4, PauliZ("q0")),
        ),
    )
    vqe = VQE(hamiltonian, algorithm_id="project.finite-shot.vqe")
    config = VQESamplingBenchmarkConfig(
        repeats=repeats,
        objective_evaluation_budget=budget,
        optimizer=OptimizerConfig(
            method="SPSA",
            max_iterations=1,
            spsa=SPSAConfig(learning_rate=0.2, perturbation=0.15),
        ),
        measurement=PauliMeasurementConfig(
            shots_per_group=shots, allocation="coefficient_l1"
        ),
        sampled_selection=SampledSelectionConfig(
            candidate_count=2, repeats_per_candidate=2
        ),
        fixed_objective_repeats=2,
        adaptive_objective_repeats=2,
        adaptive_max_objective_repeats=3,
        adaptive_standard_error_target=0.08,
        final_shots=shots,
        confidence_level=0.95,
        root_seed=29,
    )
    result = vqe.benchmark_sampling(config)
    output_dir.mkdir(parents=True, exist_ok=True)
    artifacts = {"data": str(output_dir / "benchmark.json")}
    languages: tuple[Literal["en", "zh"], ...] = ("en", "zh")
    for language in languages:
        path = output_dir / f"vqe-{language}.html"
        result.report(
            path,
            language=language,
            title="Finite-shot VQE" if language == "en" else "有限采样 VQE",
        )
        artifacts[f"report_{language}"] = str(path)
    rows = []
    for run in result.runs:
        selection = run.result.sampled_selection
        rows.append(
            {
                "strategy": run.strategy,
                "repeat_index": run.repeat_index,
                "paired_seed": run.paired_seed,
                "selected_parameters": dict(
                    run.result.selected_evaluation.parameter_bind.values
                ),
                "selected_estimator_energy": run.selected_estimator_energy,
                "exact_selected_energy": run.exact_selected_energy,
                "estimator_error": run.estimator_error,
                "paired_exact_reference_gap": run.paired_exact_reference_gap,
                "objective_evaluations": run.objective_evaluation_count,
                "objective_backend_executions": run.objective_backend_execution_count,
                "objective_shots": run.objective_shots,
                "confirmation_backend_executions": (
                    run.confirmation_backend_execution_count
                ),
                "confirmation_shots": run.confirmation_shots,
                "final_sampling_backend_executions": (
                    run.final_sampling_backend_execution_count
                ),
                "final_sampling_shots": run.final_sampling_shots,
                "diagnostic_exact_backend_executions": (
                    run.diagnostic_exact_backend_execution_count
                ),
                "selection_status": None if selection is None else selection.status,
            }
        )
    payload = {
        "sdk_version": cascaqit.__version__,
        "hamiltonian": hamiltonian.to_dict(),
        "ground_energy": 0.1 - sqrt(0.65),
        "config": config.to_dict(),
        "rows": rows,
        "benchmark": result.to_dict(),
        "artifacts": artifacts,
    }
    Path(artifacts["data"]).write_text(
        json.dumps(payload, indent=2) + "\n", encoding="utf-8"
    )
    return {
        "sdk_version": cascaqit.__version__,
        "paired_repeats": repeats,
        "objective_evaluation_ceiling": budget,
        "ground_energy": payload["ground_energy"],
        "strategies": [
            {
                "strategy": item.strategy,
                "mean_exact_selected_energy": item.mean_exact_selected_energy,
                "mean_paired_exact_reference_gap": item.mean_paired_exact_reference_gap,
                "estimator_error_rmse": item.estimator_error_rmse,
            }
            for item in result.statistics
        ],
        "total_backend_executions": sum(
            row[key]
            for row in rows
            for key in (
                "objective_backend_executions",
                "confirmation_backend_executions",
                "final_sampling_backend_executions",
                "diagnostic_exact_backend_executions",
            )
        ),
        "total_shots": sum(
            row[key]
            for row in rows
            for key in ("objective_shots", "confirmation_shots", "final_sampling_shots")
        ),
        "artifacts": artifacts,
    }


def main() -> None:
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument(
        "--output-dir", type=Path, default=Path("artifacts/finite-shot-vqe")
    )
    parser.add_argument("--repeats", type=int, default=3)
    parser.add_argument("--shots", type=int, default=128)
    parser.add_argument("--budget", type=int, default=12)
    args = parser.parse_args()
    print(
        json.dumps(
            experiment(
                args.output_dir,
                repeats=args.repeats,
                shots=args.shots,
                budget=args.budget,
            ),
            sort_keys=True,
        )
    )


if __name__ == "__main__":
    main()

下载完整脚本

连同成本一起提交比较结果

保存脚本提交号、依赖版本、JSON 和报告。比较时保留所有重复运行,包括表现较差的结果。

  1. 用解析能量公式核对每组所选参数,同时计算它与配对精确策略、与 E0 的差距。
  2. 从一份采样求值的 X、Z 计数还原能量,解释为什么末端 Z 计数本身不足以完成这一步。
  3. 执行 --shots 512 --output-dir artifacts/vqe-more-shots,比较单次求值的标准误差、实际 Backend 次数和采样成本。每次选中的能量都会改善吗?
  4. 执行 --budget 16 --output-dir artifacts/vqe-more-budget。迭代上限仍为一步,为什么只提高预算上限,实际消耗可能不变?
  5. 在副本中设置 max_iterations=3,并选择能覆盖重复求值的预算,再运行 --repeats 10。先比较配对差及其区间,再讨论推荐哪种策略。
核对思路

精确能量应不低于 E0;采样估计可能因统计误差低于它。分别用对应测量基中的 (n0−n1)/N 还原 X、Z 均值,再乘系数并加上常数。在固定状态下增加采样次数通常减少测量不确定性,但自适应停止和参数更新都可能改变整次实验,不能保证每次选中的能量都改善。没有用到的预算不会自动增加迭代。三次 SPSA 更新最多需要初始、正负扰动、最终中心共八个逻辑点:固定重复可能需要 16 次完整测量,自适应重复可能需要 24 次。预算应在观察结果之前选定。

研究扩展可以固定包含候选确认在内的总采样预算,事先规定停止规则,再用足够多的重复运行描述实际波动。讨论成本时单独列出精确诊断。这项单量子比特实验不能证明量子优势,也不能说明真实硬件性能。

English version

SDK 1.0.8a · `8b227bff`