{"task_category":"Email drafting","task_prompt":"Professional email task\n\nDraft an email to Ava, Director of Operations at TaskPrime.\n\nContext:\n- The starter pack launch needs to move from April 8 to April 13.\n- The main reason is QA found a blocking permissions issue.\n- Mention that the core work is already complete and only the final blocker remains.\n\nRequirements:\n- Keep the tone direct, calm, and professional.\n- Include a subject line.\n- Ask for confirmation on the revised date.","ideal_response":"Subject: Updated timing for the starter pack launch\n\nHi Ava,\n\nI wanted to share a quick update on the starter pack launch. We need to move the target date from April 8 to April 13 because QA found a blocking permissions issue.\n\nThe good news is that the core work is already complete, so this is about clearing the final blocker without taking unnecessary launch risk.\n\nPlease let me know if you are comfortable with the revised date of April 13. If so, I will update the plan and send the refreshed timeline this afternoon.\n\nBest,\nTaskPrime Team","difficulty_level":"medium","evaluation_criteria":["Uses a clear, relevant subject line","Covers the required facts accurately","Keeps the tone appropriate for the situation","Ends with a concrete next step"]}
{"task_category":"Email drafting","task_prompt":"Professional email task\n\nDraft an email to Priya, Product Lead at Harbor Cloud.\n\nContext:\n- The support-response workflow needs to move from June 11 to June 16.\n- The main reason is accessibility testing flagged keyboard navigation regressions.\n- Mention that the core work is already complete and only the final blocker remains.\n\nRequirements:\n- Keep the tone direct, calm, and professional.\n- Include a subject line.\n- Ask for confirmation on the revised date.","ideal_response":"Subject: Updated timing for the support-response workflow\n\nHi Priya,\n\nI wanted to share a quick update on the support-response workflow. We need to move the target date from June 11 to June 16 because accessibility testing flagged keyboard navigation regressions.\n\nThe good news is that the core work is already complete, so this is about clearing the final blocker without taking unnecessary launch risk.\n\nPlease let me know if you are comfortable with the revised date of June 16. If so, I will update the plan and send the refreshed timeline this afternoon.\n\nBest,\nTaskPrime Team","difficulty_level":"medium","evaluation_criteria":["Uses a clear, relevant subject line","Covers the required facts accurately","Keeps the tone appropriate for the situation","Ends with a concrete next step"]}
{"task_category":"Email drafting","task_prompt":"Professional email task\n\nDraft an email to Tessa, COO at SignalNest.\n\nContext:\n- You need to reschedule a meeting from 2:00 PM ET to 11:00 AM ET.\n- The meeting is about finalizing rollout responsibilities.\n- The reason for the change is that a required stakeholder was pulled into an urgent issue.\n\nRequirements:\n- Keep it polite and concise.\n- Include the original time and the new time.\n- Ask whether the new slot works.","ideal_response":"Subject: Request to move our meeting\n\nHi Tessa,\n\nI’m reaching out to see if we can move our meeting from 2:00 PM ET to 11:00 AM ET. A required stakeholder was pulled into an urgent issue, and I want to make sure the right people are in the room.\n\nThe agenda is still the same: finalizing rollout responsibilities. If the new time works for you, I’ll send an updated invite right away.\n\nApologies for the inconvenience, and thanks in advance for the flexibility.\n\nBest,\nTaskPrime Team","difficulty_level":"easy","evaluation_criteria":["Uses a clear, relevant subject line","Covers the required facts accurately","Keeps the tone appropriate for the situation","Ends with a concrete next step"]}
{"task_category":"Email drafting","task_prompt":"Professional email task\n\nDraft an email to Nina, Implementation Lead at LatticePoint.\n\nContext:\n- You need to reschedule a meeting from 10:00 AM CT to 3:00 PM CT.\n- The meeting is about aligning on training and enablement.\n- The reason for the change is that a required stakeholder was pulled into an urgent issue.\n\nRequirements:\n- Keep it polite and concise.\n- Include the original time and the new time.\n- Ask whether the new slot works.","ideal_response":"Subject: Request to move our meeting\n\nHi Nina,\n\nI’m reaching out to see if we can move our meeting from 10:00 AM CT to 3:00 PM CT. A required stakeholder was pulled into an urgent issue, and I want to make sure the right people are in the room.\n\nThe agenda is still the same: aligning on training and enablement. If the new time works for you, I’ll send an updated invite right away.\n\nApologies for the inconvenience, and thanks in advance for the flexibility.\n\nBest,\nTaskPrime Team","difficulty_level":"easy","evaluation_criteria":["Uses a clear, relevant subject line","Covers the required facts accurately","Keeps the tone appropriate for the situation","Ends with a concrete next step"]}
{"task_category":"Email drafting","task_prompt":"Sales email task\n\nDraft a cold outreach email to Tessa, COO at SignalNest.\n\nContext:\n- Pain point to reference: most AI pilots stall because the training examples are too generic\n- Product offer: a starter training pack with realistic AI task-response examples\n- Proof point to mention: every entry includes the prompt, ideal response, difficulty, and evaluation criteria\n\nRequirements:\n- Keep the email under 150 words.\n- Sound credible rather than overly salesy.\n- End with a lightweight CTA.","ideal_response":"Subject: A practical shortcut for training your AI workflows\n\nHi Tessa,\n\nI’m reaching out because most AI pilots stall because the training examples are too generic.\n\nWe built a starter training pack with realistic task-response examples so teams can move faster without starting from zero. every entry includes the prompt, ideal response, difficulty, and evaluation criteria.\n\nIf this is relevant, would you be open to a quick look at the sample preview or a short reply with what you are trying to train first?\n\nBest,\nTaskPrime Team","difficulty_level":"medium","evaluation_criteria":["Personalizes the message credibly","Explains the offer in concrete terms","Uses one light proof point instead of hype","Ends with a simple CTA"]}
{"task_category":"Email drafting","task_prompt":"Sales email task\n\nDraft a cold outreach email to Owen, Partnerships Director at Copper Lane.\n\nContext:\n- Pain point to reference: content teams often have copy, but not labeled prompt-response pairs\n- Product offer: a starter training pack with realistic AI task-response examples\n- Proof point to mention: teams use the pack for evals as well as fine-tuning\n\nRequirements:\n- Keep the email under 150 words.\n- Sound credible rather than overly salesy.\n- End with a lightweight CTA.","ideal_response":"Subject: A practical shortcut for training your AI workflows\n\nHi Owen,\n\nI’m reaching out because content teams often have copy, but not labeled prompt-response pairs.\n\nWe built a starter training pack with realistic task-response examples so teams can move faster without starting from zero. teams use the pack for evals as well as fine-tuning.\n\nIf this is relevant, would you be open to a quick look at the sample preview or a short reply with what you are trying to train first?\n\nBest,\nTaskPrime Team","difficulty_level":"medium","evaluation_criteria":["Personalizes the message credibly","Explains the offer in concrete terms","Uses one light proof point instead of hype","Ends with a simple CTA"]}
{"task_category":"Email drafting","task_prompt":"Support email task\n\nDraft a customer update email to Tessa at SignalNest.\n\nContext:\n- The reported issue was: duplicate tasks were created on double-submit\n- Resolution to share: the fix is now live in production\n- Ask the customer to retry and reply if the issue persists.\n\nRequirements:\n- Keep the tone reassuring and clear.\n- Include a subject line.\n- Avoid overly technical language.","ideal_response":"Subject: Update on the issue you reported\n\nHi Tessa,\n\nI wanted to follow up on the issue where duplicate tasks were created on double-submit. the fix is now live in production, so you can go ahead and test again when convenient.\n\nIf anything still looks off after you retry, reply here and we’ll jump back in right away.\n\nThanks for flagging this to us.\n\nBest,\nTaskPrime Support","difficulty_level":"easy","evaluation_criteria":["Explains the update clearly","Maintains an empathetic, calm tone","Includes a practical next step","Avoids unnecessary detail or filler"]}
{"task_category":"Email drafting","task_prompt":"Support email task\n\nDraft a customer update email to Lena at Atlas Freight.\n\nContext:\n- The reported issue was: mobile checklist submissions timed out on weak connections\n- Resolution to share: the request flow now retries safely on unstable networks\n- Ask the customer to retry and reply if the issue persists.\n\nRequirements:\n- Keep the tone reassuring and clear.\n- Include a subject line.\n- Avoid overly technical language.","ideal_response":"Subject: Update on the issue you reported\n\nHi Lena,\n\nI wanted to follow up on the issue where mobile checklist submissions timed out on weak connections. the request flow now retries safely on unstable networks, so you can go ahead and test again when convenient.\n\nIf anything still looks off after you retry, reply here and we’ll jump back in right away.\n\nThanks for flagging this to us.\n\nBest,\nTaskPrime Support","difficulty_level":"easy","evaluation_criteria":["Explains the update clearly","Maintains an empathetic, calm tone","Includes a practical next step","Avoids unnecessary detail or filler"]}
{"task_category":"Email drafting","task_prompt":"Support email task\n\nDraft a post-incident follow-up to Owen at Copper Lane.\n\nContext:\n- Incident: checkout redirects stalled after payment\n- Root cause to explain simply: a bad redirect environment value\n- Confirm that the issue is resolved and explain that safeguards were added.\n\nRequirements:\n- Sound accountable and calm.\n- Keep it concise.\n- Include one sentence about prevention.","ideal_response":"Subject: Follow-up on the service interruption\n\nHi Owen,\n\nI wanted to follow up after the issue where checkout redirects stalled after payment. I’m sorry for the disruption.\n\nThe incident has been resolved, and we traced it back to a bad redirect environment value. We have also added an extra safeguard in the release process so this type of issue is less likely to recur.\n\nIf you still notice anything unusual, reply here and we’ll investigate immediately.\n\nBest,\nTaskPrime Support","difficulty_level":"hard","evaluation_criteria":["Explains the update clearly","Maintains an empathetic, calm tone","Includes a practical next step","Avoids unnecessary detail or filler"]}
{"task_category":"Email drafting","task_prompt":"Support email task\n\nDraft a post-incident follow-up to Nina at LatticePoint.\n\nContext:\n- Incident: the preview page failed under peak traffic\n- Root cause to explain simply: a caching issue on the new route\n- Confirm that the issue is resolved and explain that safeguards were added.\n\nRequirements:\n- Sound accountable and calm.\n- Keep it concise.\n- Include one sentence about prevention.","ideal_response":"Subject: Follow-up on the service interruption\n\nHi Nina,\n\nI wanted to follow up after the issue where the preview page failed under peak traffic. I’m sorry for the disruption.\n\nThe incident has been resolved, and we traced it back to a caching issue on the new route. We have also added an extra safeguard in the release process so this type of issue is less likely to recur.\n\nIf you still notice anything unusual, reply here and we’ll investigate immediately.\n\nBest,\nTaskPrime Support","difficulty_level":"hard","evaluation_criteria":["Explains the update clearly","Maintains an empathetic, calm tone","Includes a practical next step","Avoids unnecessary detail or filler"]}
{"task_category":"Data analysis & summarization","task_prompt":"Review this four-week funnel for TaskPrime. Write a concise summary for the growth team. Mention the main trend, whether conversion quality improved, and one action to test next.\n\n| Week | Visitors | Signups | Activated | Paid |\n| --- | --- | --- | --- | --- |\n| W1 | 4000 | 240 | 115 | 25 |\n| W2 | 4400 | 277 | 139 | 32 |\n| W3 | 4830 | 319 | 166 | 40 |\n| W4 | 5400 | 373 | 201 | 50 |","ideal_response":"The funnel is moving in the right direction. Visitors grew from 4000 in W1 to 5400 in W4 (35.0%), and that growth carried through to paid conversions.\n\nConversion quality also improved. Activation moved from 115 out of 240 signups in W1 to 201 out of 373 in W4, and paid conversions increased from 25 to 50.\n\nThe next test should focus on the activation step, since that is still the cleanest place to unlock more paid growth from the traffic gains already happening.","difficulty_level":"medium","evaluation_criteria":["Uses the supplied numbers accurately","Comments on both volume and conversion quality","Highlights the most important movement","Ends with one practical recommendation"]}
{"task_category":"Data analysis & summarization","task_prompt":"Review this monthly revenue breakdown for TaskPrime. Summarize the performance for revenue leadership. Focus on the total trend, the strongest growth driver, and one implication for next quarter planning.\n\n| Channel | Jan | Feb | Mar |\n| --- | --- | --- | --- |\n| Organic | $6,000 | $7,200 | $8,600 |\n| Paid Search | $4,100 | $5,100 | $6,300 |\n| Referral | $3,400 | $4,300 | $5,100 |\n| Partner | $2,600 | $3,900 | $5,600 |","ideal_response":"Total revenue increased from $16,100 in January to $25,600 in March, so the quarter is trending positively overall.\n\nThe strongest growth driver was Partner, which added $3,000 over the period. That channel is doing the most to lift the total.\n\nFor next quarter, the clearest implication is to protect budget and attention around Partner while reviewing flatter channels for efficiency rather than assuming every acquisition source is equally healthy.","difficulty_level":"medium","evaluation_criteria":["Accurately compares total revenue across periods","Identifies the main growth driver correctly","Draws a reasonable planning implication","Keeps the summary concise and executive-friendly"]}
{"task_category":"Data analysis & summarization","task_prompt":"Analyze this support queue snapshot for TaskPrime. Write a short note for the support manager. Identify the busiest queue, the likely bottleneck, and one operational next step.\n\n| Queue | Tickets | First Response (hrs) | CSAT |\n| --- | --- | --- | --- |\n| Billing | 80 | 1.2 | 95% |\n| Technical | 110 | 2.7 | 91% |\n| Downloads | 55 | 1.0 | 94% |\n| Onboarding | 40 | 3.1 | 90% |","ideal_response":"Technical is the busiest queue with 110 tickets, so it is carrying the largest share of demand.\n\nThe likely operational bottleneck is Onboarding, which has the slowest first-response time at 3.1 hours. That is the queue most likely to drag down customer experience if it is not addressed.\n\nThe next step should be to review staffing coverage or workflow friction in Onboarding first, because that is the clearest path to improving service levels quickly.","difficulty_level":"medium","evaluation_criteria":["Identifies the busiest queue correctly","Uses response-time data to locate the bottleneck","Avoids inventing unsupported causes","Offers a practical operational next step"]}
{"task_category":"Data analysis & summarization","task_prompt":"Review this performance regression snapshot for TaskPrime. Summarize it for engineering. Focus on which route should be investigated first and why.\n\n| Route | Latency Before (s) | Latency After (s) | Error Before (%) | Error After (%) |\n| --- | --- | --- | --- | --- |\n| Homepage | 1.1 | 1.2 | 0.2 | 0.2 |\n| Preview | 1.4 | 2.0 | 0.3 | 0.8 |\n| Checkout | 0.9 | 1.5 | 0.1 | 0.4 |\n| Downloads | 0.8 | 1.0 | 0.2 | 0.3 |","ideal_response":"Preview should be investigated first because it now has the highest post-change latency at 2.0 seconds, making it the clearest user-facing slowdown.\n\nThere is also a meaningful reliability signal on Preview, where the error rate increased the most. That suggests the regression is not only about speed.\n\nI would start with Preview and check whether it shares a root cause with the error increase. If not, split the work quickly so both the slowest path and the noisiest failing path get attention.","difficulty_level":"hard","evaluation_criteria":["Pinpoints the route with the clearest performance regression","Considers both latency and error-rate movement","Explains the risk in plain language","Recommends a sensible technical follow-up"]}
{"task_category":"Data analysis & summarization","task_prompt":"Analyze this hiring funnel for TaskPrime. Write a short note for leadership. Call out the biggest stage drop, the overall yield from application to offer, and one recommendation.\n\n| Stage | Candidates |\n| --- | --- |\n| Applied | 320 |\n| Recruiter Screen | 88 |\n| Hiring Manager | 32 |\n| Case Study | 11 |\n| Offer | 3 |","ideal_response":"The biggest drop happens between Applied and Recruiter Screen, where the funnel loses 232 candidates. That is the stage most worth reviewing first.\n\nThe overall yield from application to offer is 0.9%, so the process is highly selective from the top of the funnel.\n\nThe best next move is to audit the criteria and speed at the Applied stage, because improvements there will have the largest downstream effect.","difficulty_level":"medium","evaluation_criteria":["Identifies the sharpest stage drop correctly","Reports the overall yield accurately","Keeps the summary useful for stakeholders","Offers a plausible process recommendation"]}
{"task_category":"Data analysis & summarization","task_prompt":"Review this four-week funnel for TaskPrime. Write a concise summary for the growth team. Mention the main trend, whether conversion quality improved, and one action to test next.\n\n| Week | Visitors | Signups | Activated | Paid |\n| --- | --- | --- | --- | --- |\n| W1 | 7200 | 432 | 207 | 46 |\n| W2 | 7600 | 479 | 240 | 55 |\n| W3 | 8030 | 530 | 276 | 66 |\n| W4 | 8600 | 593 | 320 | 80 |","ideal_response":"The funnel is moving in the right direction. Visitors grew from 7200 in W1 to 8600 in W4 (19.4%), and that growth carried through to paid conversions.\n\nConversion quality also improved. Activation moved from 207 out of 432 signups in W1 to 320 out of 593 in W4, and paid conversions increased from 46 to 80.\n\nThe next test should focus on the activation step, since that is still the cleanest place to unlock more paid growth from the traffic gains already happening.","difficulty_level":"medium","evaluation_criteria":["Uses the supplied numbers accurately","Comments on both volume and conversion quality","Highlights the most important movement","Ends with one practical recommendation"]}
{"task_category":"Data analysis & summarization","task_prompt":"Review this monthly revenue breakdown for TaskPrime. Summarize the performance for revenue leadership. Focus on the total trend, the strongest growth driver, and one implication for next quarter planning.\n\n| Channel | Jan | Feb | Mar |\n| --- | --- | --- | --- |\n| Organic | $11,000 | $12,200 | $13,600 |\n| Paid Search | $9,100 | $10,100 | $11,300 |\n| Referral | $8,400 | $9,300 | $10,100 |\n| Partner | $7,600 | $8,900 | $10,600 |","ideal_response":"Total revenue increased from $36,100 in January to $45,600 in March, so the quarter is trending positively overall.\n\nThe strongest growth driver was Partner, which added $3,000 over the period. That channel is doing the most to lift the total.\n\nFor next quarter, the clearest implication is to protect budget and attention around Partner while reviewing flatter channels for efficiency rather than assuming every acquisition source is equally healthy.","difficulty_level":"medium","evaluation_criteria":["Accurately compares total revenue across periods","Identifies the main growth driver correctly","Draws a reasonable planning implication","Keeps the summary concise and executive-friendly"]}
{"task_category":"Data analysis & summarization","task_prompt":"Analyze this support queue snapshot for TaskPrime. Write a short note for the support manager. Identify the busiest queue, the likely bottleneck, and one operational next step.\n\n| Queue | Tickets | First Response (hrs) | CSAT |\n| --- | --- | --- | --- |\n| Billing | 120 | 1.6 | 93% |\n| Technical | 170 | 2.7 | 87% |\n| Downloads | 85 | 2.0 | 94% |\n| Onboarding | 60 | 4.3 | 90% |","ideal_response":"Technical is the busiest queue with 170 tickets, so it is carrying the largest share of demand.\n\nThe likely operational bottleneck is Onboarding, which has the slowest first-response time at 4.3 hours. That is the queue most likely to drag down customer experience if it is not addressed.\n\nThe next step should be to review staffing coverage or workflow friction in Onboarding first, because that is the clearest path to improving service levels quickly.","difficulty_level":"medium","evaluation_criteria":["Identifies the busiest queue correctly","Uses response-time data to locate the bottleneck","Avoids inventing unsupported causes","Offers a practical operational next step"]}
{"task_category":"Data analysis & summarization","task_prompt":"Review this performance regression snapshot for TaskPrime. Summarize it for engineering. Focus on which route should be investigated first and why.\n\n| Route | Latency Before (s) | Latency After (s) | Error Before (%) | Error After (%) |\n| --- | --- | --- | --- | --- |\n| Homepage | 1.4 | 1.6 | 0.2 | 0.2 |\n| Preview | 1.9 | 2.6 | 0.3 | 0.9 |\n| Checkout | 1.1 | 1.9 | 0.1 | 0.4 |\n| Downloads | 1.0 | 1.3 | 0.2 | 0.3 |","ideal_response":"Preview should be investigated first because it now has the highest post-change latency at 2.6 seconds, making it the clearest user-facing slowdown.\n\nThere is also a meaningful reliability signal on Preview, where the error rate increased the most. That suggests the regression is not only about speed.\n\nI would start with Preview and check whether it shares a root cause with the error increase. If not, split the work quickly so both the slowest path and the noisiest failing path get attention.","difficulty_level":"hard","evaluation_criteria":["Pinpoints the route with the clearest performance regression","Considers both latency and error-rate movement","Explains the risk in plain language","Recommends a sensible technical follow-up"]}
{"task_category":"Data analysis & summarization","task_prompt":"Analyze this hiring funnel for TaskPrime. Write a short note for leadership. Call out the biggest stage drop, the overall yield from application to offer, and one recommendation.\n\n| Stage | Candidates |\n| --- | --- |\n| Applied | 500 |\n| Recruiter Screen | 148 |\n| Hiring Manager | 52 |\n| Case Study | 15 |\n| Offer | 4 |","ideal_response":"The biggest drop happens between Applied and Recruiter Screen, where the funnel loses 352 candidates. That is the stage most worth reviewing first.\n\nThe overall yield from application to offer is 0.8%, so the process is highly selective from the top of the funnel.\n\nThe best next move is to audit the criteria and speed at the Applied stage, because improvements there will have the largest downstream effect.","difficulty_level":"medium","evaluation_criteria":["Identifies the sharpest stage drop correctly","Reports the overall yield accurately","Keeps the summary useful for stakeholders","Offers a plausible process recommendation"]}
{"task_category":"Code generation","task_prompt":"Write a Python function called `group_invoice_totals` that groups a list of dictionaries by a string key and totals a numeric field. Skip rows with missing values, preserve two-decimal precision, and include a short usage example.","ideal_response":"```python\nfrom collections import defaultdict\nfrom decimal import Decimal, InvalidOperation\n\ndef group_invoice_totals(rows, key_name='group', value_name='amount'):\n    totals = defaultdict(Decimal)\n    for row in rows:\n        group = row.get(key_name)\n        raw_value = row.get(value_name)\n        if group in (None, '') or raw_value in (None, ''):\n            continue\n        try:\n            totals[group] += Decimal(str(raw_value))\n        except (InvalidOperation, TypeError, ValueError):\n            continue\n    return {key: value.quantize(Decimal('0.01')) for key, value in totals.items()}\n\n# Example\nprint(group_invoice_totals([{'group': 'starter', 'amount': '12.50'}, {'group': 'starter', 'amount': 7}, {'group': 'pro', 'amount': '4.25'}]))\n```","difficulty_level":"medium","evaluation_criteria":["Implements the requested behavior correctly","Handles edge cases reasonably","Uses clean Python idioms","Includes a small usage example"]}
{"task_category":"Code generation","task_prompt":"Write a Python function called `normalize_emails` that normalizes a list of strings, removes empties, deduplicates while preserving order, and includes a short usage example.","ideal_response":"```python\ndef normalize_emails(items):\n    seen = set()\n    result = []\n    for item in items:\n        if item is None:\n            continue\n        value = str(item).strip().lower()\n        if not value or value in seen:\n            continue\n        seen.add(value)\n        result.append(value)\n    return result\n\n# Example\nprint(normalize_emails(['  Example  ', 'example', None, 'Second']))\n```","difficulty_level":"easy","evaluation_criteria":["Normalizes values correctly","Removes duplicates while preserving order","Avoids unnecessary complexity","Includes a usage example"]}
{"task_category":"Code generation","task_prompt":"Write a Python function called `chunk_orders` that splits a list into chunks of a specified size, raises ValueError for invalid sizes, and includes a short usage example.","ideal_response":"```python\ndef chunk_orders(items, size):\n    if size < 1:\n        raise ValueError('size must be at least 1')\n    return [items[index:index + size] for index in range(0, len(items), size)]\n\n# Example\nprint(chunk_orders([1, 2, 3, 4, 5], 2))\n```","difficulty_level":"easy","evaluation_criteria":["Splits the list into correctly sized chunks","Validates the chunk size","Returns predictable output","Includes a usage example"]}
{"task_category":"Code generation","task_prompt":"Write a Python function called `filter_csv_by_status` that accepts a CSV string, returns matching rows as dictionaries using the standard library, and includes a short usage example.","ideal_response":"```python\nimport csv\nfrom io import StringIO\n\ndef filter_csv_by_status(csv_text, field_name, expected_value):\n    reader = csv.DictReader(StringIO(csv_text))\n    return [row for row in reader if row.get(field_name) == expected_value]\n\n# Example\nsample = 'name,status\\nAva,active\\nMilo,inactive\\nNina,active\\n'\nprint(filter_csv_by_status(sample, 'status', 'active'))\n```","difficulty_level":"medium","evaluation_criteria":["Parses CSV data correctly","Filters rows using the requested rule","Returns a clean structure","Includes a usage example"]}
{"task_category":"Code generation","task_prompt":"Write a JavaScript function called `debounceSearch` that returns a debounced version of another function, accepts a callback and delay, preserves `this`, and includes a short usage example.","ideal_response":"```javascript\nfunction debounceSearch(callback, delay) {\n  let timeoutId;\n  return function debounced(...args) {\n    const context = this;\n    clearTimeout(timeoutId);\n    timeoutId = setTimeout(() => callback.apply(context, args), delay);\n  };\n}\n\n// Example\nconst logValue = debounceSearch((value) => console.log(value), 250);\nlogValue('first');\nlogValue('latest');\n```","difficulty_level":"medium","evaluation_criteria":["Implements debounce correctly","Preserves context and latest arguments","Keeps the function reusable","Includes a usage example"]}
{"task_category":"Code generation","task_prompt":"Write a JavaScript function called `throttleScroll` that throttles a callback, executes immediately on the first call, ignores rapid repeats, and includes a short usage example.","ideal_response":"```javascript\nfunction throttleScroll(callback, interval) {\n  let lastRun = 0;\n  return function throttled(...args) {\n    const now = Date.now();\n    if (now - lastRun < interval) return;\n    lastRun = now;\n    callback.apply(this, args);\n  };\n}\n\n// Example\nconst logScroll = throttleScroll(() => console.log('scroll'), 200);\nlogScroll();\n```","difficulty_level":"medium","evaluation_criteria":["Implements throttle correctly","Avoids calling the callback too often","Keeps the code readable","Includes a usage example"]}
{"task_category":"Code generation","task_prompt":"Write a JavaScript function called `groupByStatus` that groups an array of objects by a callback result, skips nullish keys, and includes a short usage example.","ideal_response":"```javascript\nfunction groupByStatus(items, getKey) {\n  return items.reduce((groups, item) => {\n    const key = getKey(item);\n    if (key == null) return groups;\n    if (!groups[key]) groups[key] = [];\n    groups[key].push(item);\n    return groups;\n  }, {});\n}\n\n// Example\nconsole.log(groupByStatus([{ status: 'open' }, { status: 'closed' }, { status: 'open' }], (item) => item.status));\n```","difficulty_level":"easy","evaluation_criteria":["Groups items by the returned key","Returns a plain object","Avoids mutating the input","Includes a usage example"]}
{"task_category":"Code generation","task_prompt":"Write a JavaScript async function called `fetchWithTimeout` that wraps fetch with a timeout, accepts url/options/timeoutMs, aborts on timeout, and includes a short usage example.","ideal_response":"```javascript\nasync function fetchWithTimeout(url, options = {}, timeoutMs = 5000) {\n  const controller = new AbortController();\n  const timeoutId = setTimeout(() => controller.abort(), timeoutMs);\n  try {\n    return await fetch(url, { ...options, signal: controller.signal });\n  } finally {\n    clearTimeout(timeoutId);\n  }\n}\n\n// Example\nfetchWithTimeout('/api/example', {}, 3000).then((response) => console.log(response.status));\n```","difficulty_level":"medium","evaluation_criteria":["Cancels fetch when the timeout is hit","Propagates other fetch errors normally","Cleans up timeout resources","Includes a usage example"]}
{"task_category":"Code generation","task_prompt":"Write a Python function called `retry_async_task` that retries a callable a fixed number of times and raises the last error after failure.","ideal_response":"```python\ndef retry_async_task(callback, attempts=3):\n    if attempts < 1:\n        raise ValueError('attempts must be at least 1')\n    last_error = None\n    for _ in range(attempts):\n        try:\n            return callback()\n        except Exception as error:\n            last_error = error\n    raise last_error\n\n# Example\nprint(retry_async_task(lambda: 'ok', attempts=2))\n```","difficulty_level":"medium","evaluation_criteria":["Matches the requested behavior","Handles obvious edge cases","Keeps the code readable","Includes an example"]}
{"task_category":"Code generation","task_prompt":"Write a JavaScript async function called `retryAsync` that retries an async callback a fixed number of times and throws the last error after failure.","ideal_response":"```javascript\nasync function retryAsync(callback, attempts = 3) {\n  if (attempts < 1) throw new Error('attempts must be at least 1');\n  let lastError;\n  for (let index = 0; index < attempts; index += 1) {\n    try {\n      return await callback();\n    } catch (error) {\n      lastError = error;\n    }\n  }\n  throw lastError;\n}\n\n// Example\nretryAsync(async () => 'ok', 2).then(console.log);\n```","difficulty_level":"medium","evaluation_criteria":["Matches the requested behavior","Keeps the code reusable","Uses clear implementation choices","Includes an example"]}
{"task_category":"Customer support responses","task_prompt":"Write a support reply to Ava, who says they were charged twice for a TaskPrime purchase. Confirm the duplicate charge will be refunded, give a realistic timeline, and keep the tone empathetic.","ideal_response":"Hi Ava,\n\nThanks for reaching out. I’m sorry for the confusion around the duplicate charge.\n\nI’ve confirmed that one of the charges will be refunded. You should see that reversal on the original payment method within 3-5 business days.\n\nIf you do not see it after that window, reply here and we’ll check the payment status right away.\n\nBest,\nTaskPrime Support","difficulty_level":"easy","evaluation_criteria":["Acknowledges the issue with empathy","Confirms the refund outcome clearly","Sets a realistic expectation","Keeps the tone calm and concise"]}
{"task_category":"Customer support responses","task_prompt":"Write a support reply to Milo, who is upset about a delayed shipment of printed training materials from Harbor Cloud. Explain that the carrier delay is the cause, share an updated ETA, and include a make-good gesture.","ideal_response":"Hi Milo,\n\nI’m sorry for the delay with your printed materials. The shipment was slowed down by a carrier issue in transit.\n\nThe current estimate is Monday. To help make up for the inconvenience, we’ve added a 10% credit to your account.\n\nWe’ll keep monitoring the shipment and update you again if anything changes.\n\nBest,\nTaskPrime Support","difficulty_level":"medium","evaluation_criteria":["Explains the delay clearly","Provides a concrete updated expectation","Uses an empathetic tone","Includes a practical next step or make-good"]}
{"task_category":"Customer support responses","task_prompt":"Write a support reply to Priya, who requested a new feature from Atlas Freight. Thank them, avoid overpromising, share a realistic status update, and offer a workaround.","ideal_response":"Hi Priya,\n\nThanks for taking the time to share this request. Feedback like this is genuinely helpful as we prioritize what to build next.\n\nRight now, the feature is logged for product review, but we have not committed it to a release yet. In the meantime, the best workaround is to use the current export flow and keep the file structure consistent for your team.\n\nIf you want, I can also add your use case to the request so the team has more context.\n\nBest,\nTaskPrime Support","difficulty_level":"medium","evaluation_criteria":["Thanks the customer for the suggestion","Sets a realistic expectation","Offers a current workaround","Maintains a helpful tone"]}
{"task_category":"Customer support responses","task_prompt":"Write a support reply to Jonah, who cannot find their Meridian Home download links after purchase. Explain where to get the files and what to do if the redirect was missed.","ideal_response":"Hi Jonah,\n\nHappy to help. The quickest path is to use the checkout success page, which includes direct links to the JSON and JSONL downloads.\n\nIf you missed the redirect, reply here and we can resend the direct URLs. The purchase covers both formats, so you do not need to buy again.\n\nBest,\nTaskPrime Support","difficulty_level":"easy","evaluation_criteria":["Directly addresses the customer's problem","Explains the steps clearly","Uses a helpful tone","Avoids unnecessary filler"]}
{"task_category":"Customer support responses","task_prompt":"Write a support follow-up to Lena after a Copper Lane incident where hosted downloads were temporarily unavailable. Apologize, confirm resolution, explain the cause simply, and mention one prevention step.","ideal_response":"Hi Lena,\n\nI wanted to follow up after the issue where hosted downloads were temporarily unavailable. I’m sorry for the disruption.\n\nThe issue has been resolved. We traced it back to a deployment configuration problem, and we’ve added an extra public-file smoke test to our release process so this type of issue is less likely to recur.\n\nIf you still notice anything unusual, reply here and we’ll investigate immediately.\n\nBest,\nTaskPrime Support","difficulty_level":"hard","evaluation_criteria":["Acknowledges the incident clearly","Confirms that it is resolved","Explains the cause in customer-friendly language","Mentions a concrete prevention step"]}
{"task_category":"Customer support responses","task_prompt":"Write a support reply to Tessa, who is asking whether they should start with a free sample or the paid starter pack. Explain the difference and guide them toward the best next step.","ideal_response":"Hi Tessa,\n\nThe easiest way to think about it is this: the free sample is best for reviewing the schema and example quality, while the starter pack is the full working dataset.\n\nIf you still need to confirm fit, start with the sample. If the format already matches your workflow, the starter pack is the faster next step.\n\nBest,\nTaskPrime Support","difficulty_level":"easy","evaluation_criteria":["Explains the options clearly","Gives practical guidance","Keeps the reply concise","Uses a consultative tone"]}
{"task_category":"Customer support responses","task_prompt":"Write a support reply to Rafael, who asked whether the dataset files include evaluation criteria on every example. Answer clearly and offer to point them to the preview.","ideal_response":"Hi Rafael,\n\nYes, each example includes evaluation criteria alongside the task prompt, ideal response, and difficulty label.\n\nIf you want, I can also point you to the preview section so you can see the format before downloading the sample pack.\n\nBest,\nTaskPrime Support","difficulty_level":"easy","evaluation_criteria":["Answers the question directly","Keeps the message brief","Offers a helpful follow-up","Uses a friendly tone"]}
{"task_category":"Customer support responses","task_prompt":"Write a support reply to Nina, who requested a refund outside the standard refund window. Acknowledge the request respectfully, explain the policy, and offer a reasonable alternative.","ideal_response":"Hi Nina,\n\nThanks for reaching out. I understand why you asked.\n\nAt the moment, this purchase falls outside the standard refund window, so I’m not able to reverse the charge directly. The best path I can offer right now is store credit toward a future pack or help choosing the smallest option that fits your use case.\n\nIf you want to share more about what happened, I’m happy to review the details with you.\n\nBest,\nTaskPrime Support","difficulty_level":"medium","evaluation_criteria":["Explains the policy clearly","Maintains empathy","Offers a viable alternative","Keeps the tone respectful"]}
{"task_category":"Customer support responses","task_prompt":"Write a support escalation reply to Owen, who reported a missing file after a confirmed payment. Explain what is known, what is still being checked, and commit to a concrete next update.","ideal_response":"Hi Owen,\n\nThanks for flagging this. I understand the urgency.\n\nWhat we know so far is that the payment completed successfully. What we are still checking is whether the post-payment redirect or hosted file path failed on delivery.\n\nI don’t want to guess while we are still confirming the details, but I do want to keep you informed. We’ll send you a concrete update within the next hour.\n\nBest,\nTaskPrime Support","difficulty_level":"hard","evaluation_criteria":["Separates confirmed facts from active investigation","Commits to a clear next update","Uses calm, professional language","Shows appropriate empathy"]}
{"task_category":"Customer support responses","task_prompt":"Write a support reply to Grace, who asked whether the JSON and JSONL files contain the same examples. Answer clearly and mention when each format is most useful.","ideal_response":"Hi Grace,\n\nYes, the JSON and JSONL files contain the same examples.\n\nJSON is usually easier for manual review, while JSONL is often the better input format for training pipelines and batch processing.\n\nIf you want, I can also suggest which one to start with based on your workflow.\n\nBest,\nTaskPrime Support","difficulty_level":"easy","evaluation_criteria":["Answers the question directly","Explains the practical difference between formats","Keeps the tone helpful","Stays concise"]}
{"task_category":"Content creation","task_prompt":"Write a one-paragraph blog introduction for operations leaders about why AI copilots fail without realistic workflow examples. Use a practical, editorial tone, make the pain point feel real, and end by setting up the rest of the article.","ideal_response":"AI pilots rarely fail because the team lacks ambition. For operations leaders, why AI copilots fail without realistic workflow examples stops being theoretical once the examples are too generic to survive operational review. We’ll look at the signals that separate a useful starter dataset from one that only looks finished.","difficulty_level":"easy","evaluation_criteria":["Hooks the reader quickly","Matches the requested audience and topic","Uses a clear, credible editorial tone","Transitions naturally into the article"]}
{"task_category":"Content creation","task_prompt":"Write a two-paragraph blog introduction for support managers about how labeled task-response pairs improve internal tooling. Use a practical, editorial tone, make the pain point feel real, and end by setting up the rest of the article.","ideal_response":"The first draft of most training data looks plausible right up until someone tries to use it. For support managers, how labeled task-response pairs improve internal tooling becomes urgent when the prompt library looks polished, but the underlying training rows are thin. That is usually the point where a promising pilot starts to feel unreliable instead of useful.\n\nThe harder part is not imagining what the system could do. It is deciding whether the underlying examples are specific, reviewable, and grounded in real work. What follows is a practical framework for spotting weak examples early and replacing them with rows a team can actually use.","difficulty_level":"medium","evaluation_criteria":["Hooks the reader quickly","Matches the requested audience and topic","Uses a clear, credible editorial tone","Transitions naturally into the article"]}
{"task_category":"Content creation","task_prompt":"Write a two-paragraph blog introduction for product teams about why preview sections build buyer trust. Use a practical, editorial tone, make the pain point feel real, and end by setting up the rest of the article.","ideal_response":"A lot of AI workflow frustration starts long before the model answers anything. For product teams, why preview sections build buyer trust becomes urgent when the dataset is missing the edge cases that matter once usage increases. That is usually the point where a promising pilot starts to feel unreliable instead of useful.\n\nThe harder part is not imagining what the system could do. It is deciding whether the underlying examples are specific, reviewable, and grounded in real work. The goal is to leave with a simpler checklist for deciding what belongs in a production-friendly pack.","difficulty_level":"medium","evaluation_criteria":["Hooks the reader quickly","Matches the requested audience and topic","Uses a clear, credible editorial tone","Transitions naturally into the article"]}
{"task_category":"Content creation","task_prompt":"Write a one-paragraph blog introduction for founders about why starter datasets reduce time-to-pilot. Use a practical, editorial tone, make the pain point feel real, and end by setting up the rest of the article.","ideal_response":"Strong AI systems are usually built on boring discipline, not flashy demos. For founders, why starter datasets reduce time-to-pilot stops being theoretical once nobody can explain why one output is acceptable and another is not. This article breaks down what better examples look like in practice and how to tighten them without slowing the team down.","difficulty_level":"easy","evaluation_criteria":["Hooks the reader quickly","Matches the requested audience and topic","Uses a clear, credible editorial tone","Transitions naturally into the article"]}
{"task_category":"Content creation","task_prompt":"Write a two-paragraph blog introduction for revops teams about how structured content examples become reusable training data. Use a practical, editorial tone, make the pain point feel real, and end by setting up the rest of the article.","ideal_response":"Teams often blame the model when the real problem is the example set beneath it. For revops teams, how structured content examples become reusable training data becomes urgent when buyers can only see mockups instead of the real schema. That is usually the point where a promising pilot starts to feel unreliable instead of useful.\n\nThe harder part is not imagining what the system could do. It is deciding whether the underlying examples are specific, reviewable, and grounded in real work. The rest of this piece focuses on the review habits and schema choices that make training data more reliable.","difficulty_level":"medium","evaluation_criteria":["Hooks the reader quickly","Matches the requested audience and topic","Uses a clear, credible editorial tone","Transitions naturally into the article"]}
{"task_category":"Content creation","task_prompt":"Write a concise LinkedIn post that will announce a free 50-example sample pack. Use a knowledgeable, operator-style tone, open strong, stay under 100 words, and include a light CTA when appropriate.","ideal_response":"A free sample is only useful if it shows the real work.\n\nWe put 50 actual task-response examples into a public pack so teams can inspect the schema, response quality, and evaluation criteria before they buy anything. Start with the sample if you want to validate fit before moving to the full pack. If you are building an internal AI workflow, that distinction matters.","difficulty_level":"easy","evaluation_criteria":["Fits the requested channel and tone","Uses a clear hook","Communicates one useful point cleanly","Includes an appropriate CTA without hype"]}
{"task_category":"Content creation","task_prompt":"Write a concise X post that will promote a 500-example starter pack. Use a knowledgeable, operator-style tone, open strong, stay under 100 words, and include a light CTA when appropriate.","ideal_response":"Most teams do not need a giant dataset on day one. They need a starter pack with enough breadth to test training, evals, and internal workflows without inventing every example from scratch. The 500-row starter pack is built for that first real implementation step. That is the kind of detail teams notice once they move beyond demos.","difficulty_level":"medium","evaluation_criteria":["Fits the requested channel and tone","Uses a clear hook","Communicates one useful point cleanly","Includes an appropriate CTA without hype"]}
{"task_category":"Content creation","task_prompt":"Write a concise LinkedIn post that will explain why evaluation criteria belong in every row. Use a knowledgeable, operator-style tone, open strong, stay under 100 words, and include a light CTA when appropriate.","ideal_response":"Good examples should not stop at the answer.\n\nIf a row does not explain what makes the response good, it is harder to reuse for review, evals, and training quality checks. That is why every example in the pack includes evaluation criteria. Small product choices like that compound quickly.","difficulty_level":"easy","evaluation_criteria":["Fits the requested channel and tone","Uses a clear hook","Communicates one useful point cleanly","Includes an appropriate CTA without hype"]}
{"task_category":"Content creation","task_prompt":"Write a concise X post that will share one lesson about realistic training data. Use a knowledgeable, operator-style tone, open strong, stay under 100 words, and include a light CTA when appropriate.","ideal_response":"Realistic training data is usually less glamorous than people expect. The useful rows are the ones that sound like work a team actually does, not generic prompts that only look polished in a screenshot. That lesson shows up quickly once you compare mock examples to production-style tasks. It is a practical move, not a flashy one.","difficulty_level":"medium","evaluation_criteria":["Fits the requested channel and tone","Uses a clear hook","Communicates one useful point cleanly","Includes an appropriate CTA without hype"]}
{"task_category":"Content creation","task_prompt":"Write a concise LinkedIn post that will invite teams to review the sample preview. Use a knowledgeable, operator-style tone, open strong, stay under 100 words, and include a light CTA when appropriate.","ideal_response":"Preview sections matter because buyers should not have to guess.\n\nIf you are selling AI training data, show the schema, the prompt structure, and the response quality before asking anyone to commit. Reviewing a live preview is the fastest way to spot whether the pack is credible. That is usually where credibility is won or lost.","difficulty_level":"easy","evaluation_criteria":["Fits the requested channel and tone","Uses a clear hook","Communicates one useful point cleanly","Includes an appropriate CTA without hype"]}
