Skip to content
HN On Hacker News ↗

Consumer Inference Systems | Artificial Analysis

▲ 84 points 18 comments by sys42590 1w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

1 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 353
PEAK AI % 1% · §1
Analyzed
Aug 30
backend: pangram/v3.3
Segments scanned
1 windows
avg 353 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 353 words · 1 segments analyzed

Human AI-generated
§1 Human · 1%

Intelligence and Inference Performance SummaryAverage Score (16K max context) vs. End-to-End Generation TimeiPhone 17 ProSimple average of 5 evaluations chosen to represent real-world mobile device usage: BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, MATH-500 · Context limited to 16K tokens · E2E time is the seconds taken to process a 1024-token prompt and generate a 256-token responseMost attractive quadrantPareto lineA simple average of five evaluations chosen to represent real-world mobile device usage, run against models small enough to fit on portable hardware and each measured independently by Artificial Analysis. See the methodology for further details.Total wall-clock time to process a 1,024-token prompt and generate a 256-token response. Note that this is fundamental to the hardware used and the model’s architecture, and does not include the effect of model verbosity or tendency to use more or fewer turns.Inference PerformanceEnd-to-End Generation TimeiPhone 17 ProSeconds to process a 1024-token prompt and generate a 256-token response · Lower is betterTotal wall-clock time to process a 1,024-token prompt and generate a 256-token response. Note that this is fundamental to the hardware used and the model’s architecture, and does not include the effect of model verbosity or tendency to use more or fewer turns.Model IntelligenceAverage Score (Mobile Device Benchmark Set, 16K max context)iPhone 17 ProSimple average of 5 evaluations chosen to represent real-world mobile device usage: BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, MATH-500 · Context limited to 16K tokens · Higher is betterScore at 64K max contextPerformance measurements omitted as the model did not fit on this device or exceeded the time limitA simple average of five evaluations chosen to represent real-world mobile device usage, run against models small enough to fit on portable hardware and each measured independently by Artificial Analysis. See the methodology for further details.Token EfficiencyContext Budget OverrunsGenerations that stopped at the 16K-token limit instead of finishing, across every evaluation run on the model · Lower is betterEvaluation BreakdownMobile Device Benchmark Set Evaluations (16K max context)iPhone 17 ProIntelligence evaluations measured independently by Artificial Analysis · Context limited to 16K tokens · Higher is betterBFCLTool calling (index subset)IFBenchInstruction followingAA-Omniscience AccuracyKnowledgeAA-Omniscience Non-Hallucination Rate1 - hallucination rateGPQA DiamondScientific reasoningMATH-500Quantitative reasoning