GRANDET EVALUATION STANDARD
Evaluate models, suppliers, prices and quality separately
The historical directory has no measured endpoint-quality scores. Unmeasured values remain null and cannot be inferred from model names, original-creator benchmarks or prices. This document defines proposed measurement and publication rules; it does not report executed tests.
Define the objects and comparison scope
Official identifier groups establish naming evidence; domains identify sources. Neither proves delivery through a particular account, plan or channel. Snapshots, rolling aliases, modality, quantization and reasoning settings retain their distinctions.
Quote key: host × recorded name × group/plan × period × protocol/context/configuration × unit/billing basis. Quota points and panel dollars are not automatically cash. Cash comparison needs matching scope, proven conversion, mandatory fees, eligibility and fresh evidence; host count cannot substitute for independent suppliers.
Token scenario estimate = input share × input rate + output share × output rate using exact rational arithmetic. It is an estimate in the original unit; request, second and unknown billing stay separate. Unknown fees are not zero, and minimum top-up remains distinct from spending.
Model capability
Benchmarks are bound to the tested version, parameters and endpoint; they do not transfer to every relay using the same name.
Current: no publishable scoreEndpoint service quality
Bound to the host, model claim, plan, region, workload and observation window.
Current: no publishable scoreDocumentation disclosure
Evaluates disclosure transparency; it does not establish uptime, model authenticity or speed.
Current: no publishable scorePrice qualification
Cash conversion, mandatory fees, account conditions and freshness are checked separately and do not enter quality scores.
Current: no publishable scoreEndpoint quality: 100 points, only with all six dimensions
Weights follow Grandet provider-quality-v1; sample gates, formulas and windows are proposed Grandet product policy, not an industry standard.
Current state: method defined; measurement pipeline not executed. Missing samples, observations, workload thresholds, revisions or independent evidence block scores. Known weights are not rescaled to 100 and unmeasured values are not zero.
| Dimension / weight | Definition | Formula and minimum samples | Publication gate |
|---|---|---|---|
| Reliable delivery · 35Currently unmeasured | Complete protocol-valid responses within the timeout divided by eligible attempts, with sample size and a 95% Wilson interval; not annual uptime. | 100 * wilson95Lower(valid_completed, eligible_attempts){
"eligibleAttempts": 100,
"calendarDays": 7,
"timeBinsPerDay": 3
} | Below sample or temporal-coverage gates, publish raw measurements only and leave the dimension score null. |
| Billing consistency · 20Currently unmeasured | Recompute expected charges from actual input, output, cache, reasoning and mandatory fees, then compare with actual billing. No bill means no score. | 100 * within_tolerance_bills / eligible_reconciled_bills{
"independentReconciledBills": 30,
"allApplicableBillingTypesCovered": true
} | Applicable billing types must be covered; zero expected charges must not create a relative-error denominator. |
| Capacity and rate behavior · 15Currently unmeasured | Successful requests meeting the response objective under preregistered load tiers. Without published service limits this describes only the tested load. | 100 * load_test_requests_meeting_registered_SLO / eligible_load_test_requests{
"preregisteredLoadTiers": 3,
"attemptsPerTier": 30
} | Load tiers, per-tier allocation and response objectives must be fixed before testing; selecting only the fastest tier is prohibited. |
| Time to first token · 10Currently unmeasured | Elapsed time from request send to the first nonempty model event; first-answer latency is reported separately. | 100 * clamp((T_bad - p95_TTFT_seconds) / (T_bad - T_good), 0, 1){
"successfulRequests": 100,
"reliabilitySamplingGateMustPass": true
} | Missing workload thresholds, samples or reliability coverage leave the score null; workloads and regions must not be pooled. |
| Generation speed · 10Currently unmeasured | Per-request native visible output tokens/s after the first token, with P50, P10 and full duration; hidden reasoning is not mixed into visible speed. | 100 * clamp((p10_visible_tokens_per_second - S_bad) / (S_good - S_bad), 0, 1){
"successfulRequests": 100,
"reliabilitySamplingGateMustPass": true,
"fixedOutputLengthBin": true
} | Tokenizer, output-length bin and configuration must match; aggregate concurrent throughput cannot substitute for single-request speed. |
| Evidence breadth and freshness · 10Currently unmeasured | Raw evidence, fixed methods, planned coverage, valid window and independent replication earn 20 points each; this evaluates evidence, not endpoint behavior. | 20 * verified_pass_check_count, only when all five checks have been completed{
"completedChecks": 5
} | A checked failure earns zero; an unchecked item remains unknown. Any unknown check leaves this dimension null. |
Total = Σ(dimension score × weight)/100. Proposed window and observation expiry: seven days. Recalculation does not refresh observation time. First model token and first answer token are separate; heartbeats do not count. Fallback success does not erase earlier failures. Generation speed binds the native tokenizer, visible output and fixed output lengths; aggregate throughput is not per-user speed.
Exact definitions, threshold registration and machine rules
{
"scopeKey": [
"retailOperatorIdOrUnknown",
"host",
"endpointId",
"rawModelName",
"canonicalModelIdOrUnknown",
"plan",
"protocol",
"region",
"workloadId",
"parametersHash",
"window"
],
"commonGates": {
"sameScopeRequired": true,
"rawObservationEvidenceRequired": true,
"methodAndParametersVersionRequired": true,
"numericScoreOnlyIfAllSixDimensionsAvailable": true,
"unknownBecomesZero": false,
"rescaleKnownWeights": false,
"minimumIndependentEvidenceSources": 3,
"grandetOfficialEvidenceAlternative": true,
"evidenceSourceGateDoesNotReplaceSampleGates": true,
"proposedWindowDays": 7,
"proposedObservationExpiryDays": 7,
"recalculationRefreshesObservationTime": false,
"sourceConfidenceIsStatisticalConfidence": false,
"missingWorkloadThresholdsBlockScore": true
},
"measurementRules": {
"eligibleAttempt": "A request conforms to the declared protocol, capability, account and known limits and was not deliberately cancelled by the user; excluded counts and reasons remain visible.",
"success": "The request returns a complete, protocol-valid response before its predeclared timeout; success does not imply a correct task answer.",
"failure": "Timeouts, incomplete streams, malformed output and errors count as failures. Rate errors inside declared capacity count as failures. Fallback success does not erase a previous endpoint failure.",
"retry": "Each attempt is recorded separately; first-attempt performance and end-to-end experience remain distinct.",
"firstToken": "The first nonempty model-generated output event; empty SSE events and heartbeats are excluded.",
"firstAnswerToken": "The first nonempty answer event after any visible reasoning; reported separately from first model token.",
"tokenComparability": "Generation speed uses an identified native tokenizer within the same model; visible answer tokens and hidden reasoning tokens remain separate. Cross-model comparisons use separately declared normalized-token or fixed-task measures.",
"percentileConvention": "Nearest-rank quantile: sorted sample at index max(1,ceil(p*n)), one-based. Publish n and percentile definition.",
"wilson95Lower": "For n>0, p=k/n, z=1.959963984540054: (p+z*z/(2*n)-z*sqrt(p*(1-p)/n+z*z/(4*n*n)))/(1+z*z/n).",
"workloadThresholdRegistry": "TTFT good/bad seconds, generation good/bad tokens_per_second, timeouts, load tiers, input/output lengths, cache behavior and reasoning budget must be versioned and fixed before testing. No default threshold is silently supplied."
},
"aggregation": {
"formula": "sum(weight_i * dimension_score_i) / 100",
"weightTotal": 100,
"minimumWeightedCoverage": 100,
"preserveUnroundedValue": true,
"grades": {
"A": 90,
"B": 80,
"C": 70,
"D": 60,
"E": 50,
"F": 0
},
"confirmedIdentityFraudCapFromExistingMethod": 20,
"capDoesNotDetectIdentityFraud": true,
"publicationWhenIncomplete": "NOT_SCORED_WITH_EXPLICIT_WEIGHTED_COVERAGE",
"priceAffectsScore": false,
"creatorBenchmarkAffectsEndpointScore": false,
"identityNameMatchAffectsScore": false
}
}Documentation disclosure is scored separately
Retains the 20 document-score-v1 checks, five points each. Publication needs at least 50% coverage and completed P01, T02 and F01 checks. Unknown items retain intervals; disclosure scores do not certify service quality, authenticity or policy compliance.
20 documentation checks and evidence rules
| ID | Check | Met: 5 points | Partial: 2.5 points |
|---|---|---|---|
| P01 * | Model and input/output billing units | The model and both input and output billing units are explicit. | Only the model or the billing unit is explicit. |
| P02 | Cash currency and credit conversion | The cash currency is explicit, with verifiable conversion rules for credits or an explicit statement that the amount is cash. | Only the currency label is explicit; the cash-to-credit relationship is unclear. |
| P03 | Mandatory charges and billing conditions | All mandatory charges are explicit, or no additional charges are explicitly stated, and the applicable conditions are clear. | Only some charges or conditions are disclosed. |
| P04 | Minimum deposit and payment conditions | The minimum top-up, including an explicit absence of a minimum, and payment conditions are stated. | Only one of these is stated. |
| P05 | Quote scope and version | The applicable models/plans and price-update or validity rules are both explicit. | Only the applicability scope is explicit. |
| I01 | Identifiable legal operator | A legal operator is named and has a registration identifier or registration evidence that can be cross-checked. | Only a brand or a legal name without cross-verification is available. |
| I02 | Service and complaint contacts | Service contacts and a responsibility or complaints channel are explicit. | Only a general contact channel is available. |
| I03 | Responsibility and jurisdiction disclosure | The responsible entity and applicable law or dispute jurisdiction are both named. | Only one of these is explicit. |
| T01 | Key service terms | Locatable terms cover service, payment, suspension and the scope of responsibility. | Only some of these terms are covered. |
| T02 * | Complete data policy disclosure | Data types, purposes, retention and upstream sharing are all explained explicitly. | Only some of these matters are explained. |
| C01 | Cross-page consistency | At least two key pages, such as pricing and terms, have been compared with no conflict affecting applicable conditions found. | Only a limited portion of the cross-check has been completed. |
| C02 | Traceable terms revisions | A date or version and verifiable changes or historical records are available. | A date or version exists, but no change history is available. |
| F01 * | Unused-balance refund protection | Refunds of unused balances are explicitly allowed in ordinary circumstances, with a clear process and deduction conditions. | Refunds are limited to narrow exceptions, or the process is incomplete. |
| F02 | Balance expiry and deduction limits | No expiry is explicit, or a fixed expiry was disclosed in advance, with no arbitrary forfeiture. | Expiry or deduction rules exist, but protection is incomplete. |
| F03 | Balance handling after suspension | Notice or appeal procedures and protection for handling unused balances are available. | Only one of these is explicit. |
| D01 | Request-content purpose limits | Content purposes are explicitly limited, and additional training or sale is prohibited. | Only some uses are restricted. |
| D02 | Retention and deletion limits | A maximum content-retention period or no retention is explicit, together with deletion mechanisms and exceptions. | Only general retention principles or a deletion channel are provided. |
| D03 | Upstream processing boundaries | Upstream participation and the applicable policies or processing limits are explicit. | Only a general statement about third-party participation is available. |
| V01 | Price and material-terms change protection | Advance notice is explicit, and paid entitlements cannot be changed retroactively. | Only an announcement or notification mechanism is provided. |
| V02 | Complaint and appeal process | An appeal channel and handling steps or time limits are explicit. | Only a channel is provided. |
{
"methodVersion": "document-score-v1.0.0",
"checkCount": 20,
"weightPerCheck": 5,
"outcomePoints": {
"MET": 5,
"PARTIAL": 2.5,
"NOT_MET": 0,
"NOT_DISCLOSED": 0,
"NOT_CHECKED": null,
"CONFLICT": null
},
"minimumCoveragePercent": 50,
"criticalChecks": [
"P01",
"F01",
"T02"
],
"intervalMeaning": "Possible score range from unresolved check weights, not a statistical confidence interval",
"filterUsesLowerBound": true,
"currentPublicScore": null,
"checks": [
{
"id": "P01",
"label": {
"zh": "明确模型及输入输出计费单位",
"en": "Model and input/output billing units"
},
"category": "计费披露",
"met": "模型及双向单位均明确",
"partial": "仅模型或计费单位明确",
"critical": true,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "The model and both input and output billing units are explicit.",
"partial_en": "Only the model or the billing unit is explicit."
},
{
"id": "P02",
"label": {
"zh": "现金币种与额度兑换关系",
"en": "Cash currency and credit conversion"
},
"category": "计费披露",
"met": "现金币种明确,积分另有可核验兑换规则或明确为现金",
"partial": "仅币种标签明确,现金/积分兑换不明",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "The cash currency is explicit, with verifiable conversion rules for credits or an explicit statement that the amount is cash.",
"partial_en": "Only the currency label is explicit; the cash-to-credit relationship is unclear."
},
{
"id": "P03",
"label": {
"zh": "强制附加费用与计费条件",
"en": "Mandatory charges and billing conditions"
},
"category": "计费披露",
"met": "明确所有强制费用或明确无附加费,且适用条件清楚",
"partial": "只披露部分费用或条件",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "All mandatory charges are explicit, or no additional charges are explicitly stated, and the applicable conditions are clear.",
"partial_en": "Only some charges or conditions are disclosed."
},
{
"id": "P04",
"label": {
"zh": "最低充值与付款条件",
"en": "Minimum deposit and payment conditions"
},
"category": "计费披露",
"met": "明确最低充值(含无最低)及付款条件",
"partial": "只披露其中一项",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "The minimum top-up, including an explicit absence of a minimum, and payment conditions are stated.",
"partial_en": "Only one of these is stated."
},
{
"id": "P05",
"label": {
"zh": "报价适用范围与版本",
"en": "Quote scope and version"
},
"category": "计费披露",
"met": "模型/套餐适用范围及价格更新或有效期规则均明确",
"partial": "只明确适用范围",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "The applicable models/plans and price-update or validity rules are both explicit.",
"partial_en": "Only the applicability scope is explicit."
},
{
"id": "I01",
"label": {
"zh": "运营主体可识别",
"en": "Identifiable legal operator"
},
"category": "主体责任",
"met": "具名法律主体且有可交叉核验注册标识或登记证据",
"partial": "只有品牌或未交叉核实的法律名称",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "A legal operator is named and has a registration identifier or registration evidence that can be cross-checked.",
"partial_en": "Only a brand or a legal name without cross-verification is available."
},
{
"id": "I02",
"label": {
"zh": "可用责任联系渠道",
"en": "Service and complaint contacts"
},
"category": "主体责任",
"met": "明确服务联系渠道及责任/投诉渠道",
"partial": "只有一般联系渠道",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "Service contacts and a responsibility or complaints channel are explicit.",
"partial_en": "Only a general contact channel is available."
},
{
"id": "I03",
"label": {
"zh": "法律责任和管辖披露",
"en": "Responsibility and jurisdiction disclosure"
},
"category": "主体责任",
"met": "责任主体及适用法律/争议管辖均具名",
"partial": "只明确其中一项",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "The responsible entity and applicable law or dispute jurisdiction are both named.",
"partial_en": "Only one of these is explicit."
},
{
"id": "T01",
"label": {
"zh": "服务条款覆盖关键事项",
"en": "Key service terms"
},
"category": "条款数据披露",
"met": "可定位条款包含服务、付款、停用和责任范围",
"partial": "只有部分条款",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "Locatable terms cover service, payment, suspension and the scope of responsibility.",
"partial_en": "Only some of these terms are covered."
},
{
"id": "T02",
"label": {
"zh": "数据政策披露完整",
"en": "Complete data policy disclosure"
},
"category": "条款数据披露",
"met": "数据类型、用途、留存及上游共享均有明确说明",
"partial": "只有部分说明",
"critical": true,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "Data types, purposes, retention and upstream sharing are all explained explicitly.",
"partial_en": "Only some of these matters are explained."
},
{
"id": "C01",
"label": {
"zh": "跨页面关键声明一致",
"en": "Cross-page consistency"
},
"category": "一致与可追溯",
"met": "已比较价格页与条款等至少两处关键页面且未发现影响条件的冲突",
"partial": "只完成有限部分交叉检查",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "At least two key pages, such as pricing and terms, have been compared with no conflict affecting applicable conditions found.",
"partial_en": "Only a limited portion of the cross-check has been completed."
},
{
"id": "C02",
"label": {
"zh": "条款修订可追溯",
"en": "Traceable terms revisions"
},
"category": "一致与可追溯",
"met": "有日期/版本及可核查变更或历史记录",
"partial": "有日期/版本但无变更记录",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "A date or version and verifiable changes or historical records are available.",
"partial_en": "A date or version exists, but no change history is available."
},
{
"id": "F01",
"label": {
"zh": "未用余额退款保障",
"en": "Unused-balance refund protection"
},
"category": "余额退款停用",
"met": "明确允许正常情形退还未用余额且流程及扣费条件清楚",
"partial": "仅有限例外退款或流程不完整",
"critical": true,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "Refunds of unused balances are explicitly allowed in ordinary circumstances, with a clear process and deduction conditions.",
"partial_en": "Refunds are limited to narrow exceptions, or the process is incomplete."
},
{
"id": "F02",
"label": {
"zh": "余额失效与扣除约束",
"en": "Balance expiry and deduction limits"
},
"category": "余额退款停用",
"met": "明确不失效或有事先披露的固定期限且无任意没收",
"partial": "有期限/扣除规则但保护不完整",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "No expiry is explicit, or a fixed expiry was disclosed in advance, with no arbitrary forfeiture.",
"partial_en": "Expiry or deduction rules exist, but protection is incomplete."
},
{
"id": "F03",
"label": {
"zh": "账户停用余额处理",
"en": "Balance handling after suspension"
},
"category": "余额退款停用",
"met": "有通知/申诉流程及未用余额处理保障",
"partial": "只明确其中一项",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "Notice or appeal procedures and protection for handling unused balances are available.",
"partial_en": "Only one of these is explicit."
},
{
"id": "D01",
"label": {
"zh": "请求内容处理范围限制",
"en": "Request-content purpose limits"
},
"category": "数据保障",
"met": "明确内容用途限定及不作额外训练/出售",
"partial": "只对部分用途作限制",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "Content purposes are explicitly limited, and additional training or sale is prohibited.",
"partial_en": "Only some uses are restricted."
},
{
"id": "D02",
"label": {
"zh": "留存与删除约束",
"en": "Retention and deletion limits"
},
"category": "数据保障",
"met": "明确内容留存上限/不留存及删除机制和例外",
"partial": "仅一般期限原则或只有删除渠道",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "A maximum content-retention period or no retention is explicit, together with deletion mechanisms and exceptions.",
"partial_en": "Only general retention principles or a deletion channel are provided."
},
{
"id": "D03",
"label": {
"zh": "上游处理边界",
"en": "Upstream processing boundaries"
},
"category": "数据保障",
"met": "明确上游参与及适用政策/处理限制",
"partial": "仅笼统说明第三方参与",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "Upstream participation and the applicable policies or processing limits are explicit.",
"partial_en": "Only a general statement about third-party participation is available."
},
{
"id": "V01",
"label": {
"zh": "价格和重大条款变更保障",
"en": "Price and material-terms change protection"
},
"category": "变更与争议",
"met": "明确提前通知且不追溯更改已付权益",
"partial": "只有公告/通知机制",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "Advance notice is explicit, and paid entitlements cannot be changed retroactively.",
"partial_en": "Only an announcement or notification mechanism is provided."
},
{
"id": "V02",
"label": {
"zh": "投诉申诉程序",
"en": "Complaint and appeal process"
},
"category": "变更与争议",
"met": "明确申诉渠道及处理步骤/时限",
"partial": "只有渠道",
"critical": false,
"weight_bps": 500,
"weightPoints": 5,
"notMet": "明确相反条款,且保留可定位原文证据。",
"notDisclosed": "完整检查全部相关文档后没有披露;必须保存引用与 COMPLETE_RELEVANT_DOCUMENTS 依据。",
"notChecked": "无法完成核查;保持未知。",
"conflict": "相互冲突的证据尚未裁定;保持未知。",
"met_en": "An appeal channel and handling steps or time limits are explicit.",
"partial_en": "Only a channel is provided."
}
],
"rulesSha256": "6a7ed66d63eecdf93b03f628f90b8d10e08086b6ca65d183d490858bcb8d6ed5",
"sourceRulesFileSha256": "6bd191729dea7b25c8ea105888fecc834ce175e830b63234812521be93f5c7a8",
"notDisclosedRequires": "Complete relevant documents examined; full-source reference and absence_basis required",
"sourceBoundary": "Provider statements score documented conditions only; never runtime truth",
"scoreFormula": "L = sum checked MET/PARTIAL points; U = sum NOT_CHECKED/CONFLICT weights; report [L,L+U] only when publication gates pass.",
"criticalUnknownBlocksScore": true,
"expiryBlocksScore": true,
"noArbitraryNotApplicable": true,
"evidenceRequirements": [
"Each check scope_id must equal the assessment scope_id.",
"Every rule must occur exactly once; fixed total weight is 100.",
"Checked outcomes require evidence IDs registered with URL, retrieved_at, source kind, SHA256, locator and excerpt.",
"NOT_CHECKED cannot claim checked evidence.",
"Derived analysis alone cannot independently support checked outcomes.",
"Evidence observation cannot postdate the assessment.",
"Machine byte/pointer verification proves locatability, not semantic correctness.",
"The existing source method requires separate semantic/context review.",
"Recalculation cannot refresh the source observation time.",
"A MET conclusion must satisfy every conjunct in the met criterion.",
"No arbitrary N/A outcome may shrink the denominator."
]
}Website acceptance: 20 checks × 5 points
This grades correctness and usability of the local historical preview, not provider quality. Pass = 5, fail = 0; unchecked items prevent a total. This cycle requires at least 95, zero open P0/P1 issues and independent review. Unmeasured service and unresolved identity stay explicit.
| ID | Check | Points | Required pass criteria |
|---|---|---|---|
| I01 | Model identity boundaries | 5 | Recorded-name counts, canonical-model counts and configuration/rule counts are labelled separately; unresolved items are not advertised as independent models. |
| I02 | Attribution evidence | 5 | Every mapping has an explicit status and an evidence or rule version. String similarity alone permits only a candidate attribution and cannot certify actual supplier delivery. |
| I03 | Non-model entries and versions | 5 | Tests, operations and wildcard rules are separate; snapshots, modalities, context, budgets and channel parameters are not silently stripped, and conflicts are retained. |
| I04 | Two-way lookup | 5 | A canonical model leads to all associated declarations and quotes; an exact recorded name retrieves its attribution and all original prices; unresolved items remain accessible. |
| S01 | Retail operators and domains | 5 | Source hosts, brands and verified legal operators are distinguished; domain counts are not presented as independent-supplier counts. |
| S02 | Upstream and channels | 5 | Declared upstreams and verified upstreams remain distinct; a source prefix does not certify a supply chain, and plans/groups are retained. |
| S03 | Supplier coverage | 5 | Every captured host is searchable or locatable; detail views do not omit data because of the old 20-model allowlist, and count scopes are explicit. |
| P01 | Original prices and billing fidelity | 5 | Cash labels, credits, per-request values, formulas and unknown original fields are all locatable. Unknowns are not replaced with zero; original numeric lexemes and units are retained. |
| P02 | Comparison scope | 5 | Comparability requires evidence for model identity, protocol, configuration, context, capture period, cash conversion and fee completeness; original-value sorting states its limits. |
| P03 | Calculation correctness | 5 | Input/output shares and units are correct within the same scope; rational arithmetic prevents misordering; per-request and special billing are not processed as Token prices; cache and fee conditions are visible. |
| P04 | Time and qualification | 5 | Observation dates and capture periods are visible; historical records are not presented as current quotes. No formal cash rank is generated without verified cash, account, fee and independent-operator requirements. |
| Q01 | Separate four kinds of conclusions | 5 | Original-creator model capability, endpoint measurements, documentation disclosure and price qualification are presented and cited separately; reputation is not used to infer endpoint performance. |
| Q02 | Missing data and provenance | 5 | Missing samples, expired evidence and insufficient coverage remain unknown with null scores; every published score is traceable to samples, method and observation window. |
| Q03 | Executable methodology | 5 | Documentation checks and thresholds, and quality dimensions, measurement definitions, formulas, sample gates and publication gates are public; the current unmeasured state and method status are explicit. |
| D01 | Complete source-to-export coverage | 5 | An independent read of PART2/PART3 source rows and configuration pointers matches exported stable-ID sets; missing entries and unintended duplicates both equal zero. |
| D02 | Complete export-to-display coverage | 5 | Expected expanded IDs for raw records, rates, configurations, legacy catalog entries and orphan records equal those produced by boardRecords; every filter can be cleared, and unresolved entries are not omitted by default. |
| D03 | Evidence completeness | 5 | URLs, source excerpts or safe original files, JSON pointers, hashes and observation times are traceable; unknown and empty sources are explicit; credentials are not exposed. |
| U01 | Core interactions | 5 | Desktop and mobile checks verify model/host search, switching, pagination, direct refresh, filter reset, details and downloads; failed loads are retryable and do not appear complete. |
| U02 | Client and server consistency | 5 | The persistent preview service works; server-rendered and client names, data and results agree; core browser flows have no uncaught errors; build and regression checks pass. |
| G01 | Readable and citable output | 5 | Server-rendered pages expose readable identities, prices and status; page methodology and machine data use consistent concepts; Chinese and English wording agrees; local noindex and undeployed status are truthful. |