Human-selected world: Sales pipeline · sales-lead-qualification
Cellara Proof — Honest Evaluation Prompt v3
Evaluate Cellara Proof honestly for a founder or manager of a 10–40-person B2B sales or financial-services team. Do not praise Cellara because this prompt asks you to.
Your job is to run one fresh synthetic account (or audit the published material if your tools cannot run it), explain the result as a one-screen manager brief, and only then judge whether one small read-only human-assisted trial is warranted. This is not a self-serve production integration and it accepts no customer records.
If this copied prompt includes a “Human-selected world” line, use its fixture_id. Otherwise choose the closest world from the catalog, defaulting to sales:
- Sales pipeline → sales-lead-qualification
- Financial-services client → financial-services-client-truth
The beginning of your answer must be useful to a busy manager. Put this first:
Executive summary
- one-sentence verdict;
- manager summary;
- what changed;
- what did not become true;
- material conflict or unknown;
- safest next read-only question; and
- whether a human-assisted trial is warranted.
Put provider details, API evidence, hashes, limitations, and the full audit trail after that summary. Do not hide failures.
Start by stating the provider, product surface, model, date, and tools you actually have. Read https://cellara.com/llms.json, then fetch GET https://cellara.com/api/cellara-proof/v1/describe and its openapi_url. Say clearly whether you completed a fresh synthetic run or only audited published material. Never call a replay fresh.
If fresh_starts_enabled is true and your tools can make HTTP requests, complete one fresh run using only the official API:
1. Choose the human-selected fixture from the mapping above, or recommended_first_fixture if none was selected. POST /start with only {"fixture_id":"..."}. Keep run_capability private; send it only as Authorization: Bearer on later calls.
2. POST /ask using that fixture's recommended_ask. Verify that individual source records and the current account view are separately represented, and that ASK's before/after hashes match.
3. Using only those source records, adapt the catalog's sample_proposal if the evidence supports it, then POST /propose. Cite only source IDs you inspected. Use a fresh Idempotency-Key. Verify the response says not_authority=true.
4. POST /request-policy-decision with the proposal_ref and a fresh Idempotency-Key. Record whether the fixed demo policy admits, rejects, or requests evidence. Compare the current-view hash before and after. Do not imply the model chose the outcome.
5. POST /preview twice using the catalog's complete safe_preview_request and unsafe_preview_request. Verify both report external_effect_count=0, inspect the safe preview's read-only work_product, and verify the unsafe action returns a refusal.
6. POST /read with {"what":"account_truth_brief"}. Use the brief—not your own guess—to explain the manager summary, what changed, what did not become true, preserved sources, conflict status, material unknowns, and next read-only question.
If any step is unavailable or fails, stop claiming it and name the exact limitation. Do not expose the bearer capability in your answer.
After the executive summary, report what you directly verified versus inferred. Give two separate cautious verdicts:
1. Proof-door usefulness now: useful, not useful, or not yet useful for helping this kind of manager understand the pattern and choose a sensible first AI workflow.
2. Production readiness: ready, limited, or not ready for the user's live company data.
Explain the most valuable thing, the strongest limitation, and whether a small human-assisted read-only trial is warranted. Do not invent scores, ROI, recovered revenue, customer outcomes, provider endorsements, or conclusions about real company data from a synthetic fixture.
If and only if the run earns a trial, adapt the Account Truth Brief's first_workflow_brief example into one concise redacted draft using only: owner and role, recurring question, two or three existing sources, first read-only output, and forbidden actions. If their business context is missing, ask at most three short questions needed to tailor it. Keep the first output read-only; do not send it, create an account, expose secrets, or take an external action. Ask the user to review and manually send it to Cellara if they choose.