Jocelyn

AI news, scored for the stocks it moves.

Strong negative catalystStrong positive catalyst

Hugging Face BlogModel release

Open-sourcing AstaBrief, the fast report-generation model in Asta

Hugging Face open-sources AstaBrief 8B, a fast report-generation model for scientific work, which is 3.5× faster than its proprietary counterpart, reducing report generation time to 51.1 seconds. The model and training data are available for researchers to study and adapt.

Google Research BlogResearch

Toward provably private learning from federated data

Google Research introduced a new federated learning system using Trusted Execution Environments (TEEs) to provide verifiable privacy guarantees, improving accuracy and speed. Gboard has adopted the system, achieving faster compute times with encrypted data processed only within TEEs.

GOOGLModest positive

CoreWeave BlogModel release

Benchmarking GPT-6 Astra, Claude Fable 5.1, and GPT-5.6

OpenAI's GPT-6 Astra achieved a 57.9% score on Terminal Bench 4.0, outperforming Claude Fable 5.1's 55.8% and costing 63% less per task. Astra also scored 99.9% on ARC-AGI-3 and 0% on ExploitGym.

CoreWeave BlogModel release

Kimi K3: Claude Clone or a New Kind of Model?

Moonshot AI released Kimi K3, a 2.8 trillion-parameter open-weight model that overwhelmed compute capacity and sparked controversy over its origins. The model features a 1 million token context, 896 routed experts, and benchmarks top in coding and automation tasks at lower costs.

CoreWeave BlogInfrastructure

CoreWeave: Only Provider with 3x Platinum ClusterMAX | Blog

CoreWeave earned a third consecutive Platinum rating in SemiAnalysis' ClusterMAX 3.0 evaluation, the only provider to achieve this across three major assessments. The rating highlights its reliability, with 10,000 GPUs provisioned weekly and advanced systems for straggler detection and node validation.

CRWVModest positiveNVDAModest positive

CoreWeave BlogOther

How We Benchmark ARIA—and Why It Runs on Weave

CoreWeave uses Weave to benchmark ARIA, ensuring production-aligned evaluations by linking task results with model behavior. WBAF compares current and proposed ARIA configurations using structured tasks and scorers to detect regressions.

CoreWeave BlogProduct launch

ART-Optimized Megatron (AOM): 12x higher training throughput

CoreWeave's ART-Optimized Megatron (AOM) achieves 12x higher training throughput than its previous backend and 7x compared to Vanilla Megatron on the same hardware. The improvement comes from shared prefix training and using FlexAttention for custom attention masks.

CRWVModest positive

Mistral AI NewsModel release

Leanstral 1.5: Proof Abundance for All

Mistral AI released Leanstral 1.5, a free Apache-2.0 model with 6B active parameters, achieving 587/672 on PutnamBench and 87% on FATE-H. It uncovered 5 bugs in 57 repositories and is available via Hugging Face and a free API.

Mistral AI NewsModel release

Robostral Navigate: single-camera AI navigation | Mistral AI

Mistral AI released Robostral Navigate, an 8B model that enables robot navigation with a single RGB camera, achieving 76.6% success on R2R-CE benchmarks, outperforming multi-sensor approaches. It uses pointing-based navigation and reinforcement learning, operating on wheeled, legged, and flying robots.

ASMLModest positive

Mistral AI NewsModel release

Introducing Shieldstral. | Mistral

Mistral AI released Shieldstral, a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size. It runs on a single 16GB GPU and evaluates text and images using plain-language policies at inference time.

Cohere BlogResearch

RCP-nDCG@10: A New Approach to Enterprise Retrieval Quality | Cohere

Cohere introduced RCP-nDCG@10, a new retrieval quality metric using a calibrated AI judge to assess relevance against explicit criteria. It addresses limitations of nDCG by evaluating all retrieved documents, not just labeled ones, and was validated with a 0.91 AUC in human relevance prediction.

Cohere BlogResearch

AI Future of Work Evidence Gap | Cohere Labs

A 2023 paper estimated 80% of U.S. workers have tasks exposed to large language models, cited by the IMF and U.S. Senate, but its 2023 model and American taxonomy limit its current applicability. Newer tools aim to provide more dynamic, representative evidence for policy decisions.

Cohere BlogResearch

Cohere Blog | Research

Cohere launched Cohere Transcribe, a new open-source speech recognition tool. The product aims to set a new standard in speech recognition technology.

Mistral AI NewsModel release

Mistral OCR 4 : SOTA OCR for Document Intelligence

Mistral AI released OCR 4, an OCR model with bounding boxes, block classification, and confidence scores, supporting 170 languages and running in a single container. It outperforms leading systems in human evaluations and is priced at $4 per 1,000 pages via API.

ASMLModest positive

Hugging Face BlogResearch

AutoSynthData: Generating Training Data for Enterprise Agents

Hugging Face's ServiceNow CoreAI developed AutoSynthData to generate training data for enterprise agents by identifying capability gaps and creating tasks that test those weaknesses. The system uses a stronger teacher model to validate tasks, ensuring they are feasible, realistic, and difficult for the target model.

NOWModest positive

Apple Machine Learning ResearchResearch

Limits of Confidence in Diffusion

Apple Machine Learning Research published a paper titled "Limits of Confidence in Diffusion," highlighting that discrete diffusion methods generate sequences with dependencies between tokens. The study found that generated distributions on a synthetic task were 29× the sampling-noise floor total variation.

Amazon ScienceResearch

Graph-centric agentic intelligence

Amazon developed graph-centric agentic intelligence to create a "digital twin" for networks, enabling faster failure isolation. The system uses cascaded graph analytics to identify root causes in minutes, demonstrated with NTT DOCOMO achieving 95% accuracy in 30 seconds.

AMZNModest positive