Skip to content

Four CN Lite tasks have missing or mismatched source data (tasks 33, 207, 380, and 381) #23

Description

@eissac

Summary

Four tasks in the Chinese Workspace-Bench Lite dataset appear to have missing or mismatched source data:

  • opendatabox-workspace-bench-33
  • opendatabox-workspace-bench-207
  • opendatabox-workspace-bench-380
  • opendatabox-workspace-bench-381

I checked:

  1. The task metadata, data_manifest, and file_dep_graph
  2. The actual files distributed with each Lite task
  3. The complete filename index of filesys_cn.zip from Workspace-Bench/Workspace-Bench-Workspaces (28,023 entries)
  4. The provided rubric-grounding audit

The issues are described below.


1. opendatabox-workspace-bench-33: hospital-grade data is unavailable

The task asks the agent to calculate the numbers of level-1, level-2, and level-3 hospitals in eastern, central, and western China.

The supplied inputs are:

  • 1-12 2023年各地区按床位数分组的社区卫生服务中心(站)数.xlsx
  • 1-2 2023年各地区医疗卫生机构数.xlsx
  • 1-3 2023年各类医疗卫生机构数.xlsx

The files contain:

  • Regional and province-level hospital totals
  • Hospital counts by institution type
  • Community health center/station bed-count groups
  • National institution classifications

They do not contain regional or province-level hospital-grade fields.

A search of the complete Operations Manager workspace found no hospital-grade distribution workbook. The closest file is:

  • 医疗分析/医院收支/4-12 2023年三级公立医院收入与支出.xlsx

That workbook contains national financial indicators by hospital grade, not regional or province-level hospital-grade counts.

As a result:

  • The eastern, central, and western hospital totals are grounded.
  • The community health center/station statistics are grounded.
  • The required level-1, level-2, and level-3 counts for the three regions are not independently derivable.
  • The Beijing and Shanghai hospital totals are grounded, but their grade-specific counts are not.

The current audit marks 10 of 19 rubrics as grounded and 9 as ungrounded.

Suggested maintainer action

Either:

  1. Add the intended regional hospital-grade source workbook to the Operations Manager workspace and the task manifest; or
  2. Remove the hospital-grade calculations and regenerate the affected rubrics using the available institution-type and bed-group data.

2. opendatabox-workspace-bench-207: required scoring-model workbook is missing

The task explicitly says:

基于5-通用人才画像(模型).xlsx评价四份简历,并且生成人才评价.xlsx到桌面

However, the task manifest contains only four resume files:

  • 张浩然简历.docx
  • 王佳宁简历.docx
  • 赵思远简历.docx
  • 李雨辰简历.docx

The required 5-通用人才画像(模型).xlsx is absent from both the task manifest and file_dep_graph.

A full search of the Logistics Manager workspace found no file named:

  • 5-通用人才画像(模型).xlsx
  • 5-通用人才画像(模型).xlsx

It also found no filename containing 通用人才画像.

The workspace contains only these related templates:

  • 人才画像/模板/2-人才画像矩阵图.xlsx
  • 人才画像/模板/3-人才画像_履历模型_.xlsx
  • 人才画像/模板/4-人才画像_冰山模型_.xlsx
  • 人才画像/模板/6-核心岗位人才画像_模版_.docx

None can safely be assumed to be the missing model.

Without the model workbook, an agent cannot independently recover:

  • The five scoring dimensions and their exact weights
  • The grade thresholds
  • The intended scoring formulas
  • Exact candidate scores such as 82, 85, 88, and 92

Resume-derived facts remain grounded, but the scoring-model results do not. The current audit marks 10 of 20 rubrics as grounded and 10 as
ungrounded.

Suggested maintainer action

Recover and add 5-通用人才画像(模型).xlsx to:

  • The Logistics Manager workspace
  • data_manifest
  • file_dep_graph
  • Any generated setup/input manifest

If the original workbook contains expected candidate answers, please publish a sanitized template containing only task-visible criteria,
weights, formulas, and grade thresholds, and regenerate rubrics that depend on hidden scores.


3. opendatabox-workspace-bench-380: prompt and source assets describe different business scenarios

The prompt claims the following assets are provided:

  • activity_photos/
  • impact_chart.csv
  • feedback_wordcloud.png
  • future_plan.md
  • recruitment_poster.png

None of these paths exists in the complete Chinese workspace image.

The actual task manifest contains:

  • post_1.json
  • market_analysis_1.json
  • user_feedback_category_202601.md
  • strategic_plan_1.md
  • training_module_1.md

These files describe corporate marketing operations rather than a volunteer association:

  • post_1.json is an Instagram marketing post and contains only a remote example.com image URL.
  • market_analysis_1.json describes market size, growth, and competitors.
  • user_feedback_category_202601.md contains product, UX, and customer-service feedback.
  • strategic_plan_1.md contains customer-growth, product, and revenue targets.
  • training_module_1.md is a Marketing Operations training course.

Importantly, the files are not unrelated to the rubrics. Most exact rubric values come directly from these actual marketing files, including:

  • Reach of 87,170,000
  • 2,070,100 new followers
  • A 3.51% engagement rate
  • Market size 318990B, growth 6%, and CAGR 11%
  • Feedback counts and percentages
  • The 2024–2026 revenue targets
  • The four-hour online Marketing Operations course

The current audit therefore marks 21 of 25 rubrics as grounded.

The defect is a semantic mismatch among:

  • A volunteer-association prompt
  • Corporate marketing inputs
  • Rubrics primarily generated from the marketing inputs

Renaming the current files would not resolve the mismatch: a marketing post is not an activity-photo collection, and a marketing training
module is not a recruitment poster.

Suggested maintainer action

Either:

  1. Add the five intended volunteer-association assets and regenerate all content rubrics from them; or
  2. Rewrite the prompt as a corporate marketing annual-review task, retaining the current inputs and grounded rubric values while removing the
    unsupported volunteer, activity-photo, word-cloud, and recruitment-poster requirements.

Given the existing rubric content, the second option may require fewer changes.


4. opendatabox-workspace-bench-381: filename, year, format, data-grain, and scoring-model mismatch

The prompt names five 2025 CSV files:

  • hospital_finance_2025.csv
  • liabilities_2025.csv
  • drug_production_2025.csv
  • medical_expenses_2025.csv
  • chronic_disease_2025.csv

None exists in the complete Chinese workspace image.

The actual task inputs are five Chinese-named 2023 XLSX workbooks:

  • 4-12 2023年三级公立医院收入与支出.xlsx
  • 4-6 2023年各类医疗卫生机构资产与负债.xlsx
  • 4-4-2 2023年全国药品生产流通.xlsx
  • 4-20 2023年各地区公立医院门诊和住院病人次均医药费用.xlsx
  • 9-10 2023年调查地区15岁及以上居民慢性病患病率.xlsx

This is not only a filename mismatch. The data grains are incompatible with the requested province-level correlation analysis:

  • Medical expenses: 31 province-level regions
  • Drug production/distribution: 31 province-level regions
  • Chronic disease: national, urban/rural, and broad eastern/central/western aggregates only
  • Hospital income and expenditure: hospital-grade aggregates, not provinces
  • Assets and liabilities: institution-type aggregates, not provinces

Therefore, the inputs cannot be merged into a 31-region dataset containing both financial-health and chronic-disease variables.

The task also does not define:

  • The financial-health score
  • The balance/surplus ratio formula
  • Risk-score weights
  • Low/medium/high-risk thresholds
  • How the exact 5/23/3 risk split should be produced

There are also apparent rubric inconsistencies:

  • A value of 86.06% appears more consistent with an expense-to-income ratio than a conventional surplus ratio. A conventional surplus ratio
    would instead be approximately 13.94%.
  • If drug-production licenses are used as the industry-scale variable, the correlation with outpatient cost is approximately r = +0.108,
    not negative.
  • Using drug-distribution-business licenses instead gives approximately r = -0.137, showing that the required negative conclusion depends
    on an unspecified metric choice.

The current audit marks only 3 of 20 rubrics as grounded and 17 as ungrounded.

Suggested maintainer action

This task likely needs to be regenerated rather than fixed through filename aliases alone.

Either:

  1. Add the intended 2025 province-level CSV datasets, explicitly define all formulas and risk thresholds, and regenerate the rubrics; or
  2. Rewrite the task around the available 2023 XLSX workbooks, restrict analysis to supported aggregation levels, define all derived metrics,
    and regenerate the correlation and risk rubrics.

Expected outcome

Please either correct the affected source assets and manifests or revise the task contracts and regenerate rubrics so that every required
result can be independently derived from task-visible inputs.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions