Daily Digest — 2026-09-13
7 items · 5 research labs, 2 industry media
🏛️ Research Labs (5)
Perplexity trusts GPT-6 Astra with end-to-end systems
Perplexity employs GPT-6 Astra for end-to-end system automation, including code generation, communication drafting, and production monitoring, with reduced human oversight compared to prior models. The model's improved code-writing capability enhances Perplexity's search engine by synthesizing web and internal data concisely. GPT-6 Astra generates realistic API responses for testing workflows, enabling comprehensive application validation without manual intervention, as reported by Perplexity's Chief Strategy Officer.
gpt-6 astraend-to-end systemscode generationapi testingproduction monitoring
Cognition helps Devin test its own work with GPT‑6 Astra
Cognition integrates GPT-6 Astra to enhance Devin, an autonomous software engineer, by improving its ability to test and validate code functionality. GPT-6 Astra enables Devin to generate test recordings and reports, such as simulating iPhone game Otter Run and identifying untested areas, reducing manual code review effort. Additionally, Devin leverages Astra to rapidly address customer-reported bugs by fixing issues and returning visual evidence. Cognition anticipates this integration will streamline code review processes, allowing engineers to inspect less code manually while increasing deployment efficiency. The approach is applied across Cognition’s product lineup, including Devin’s core cloud agent, CLI, and desktop products.
autonomous software engineercode reviewtest recordingsbug fixingcloud agent
Gemini API Managed Agents: 3.6 Flash, hooks, and more
The Gemini API introduces Managed Agents 3.6 Flash with environment hooks, model selection, and free-tier access, enhancing autonomous cloud sandbox execution for coding tasks. Key features include pre/post-tool execution hooks (via regex-matched commands or HTTP endpoints), budget controls (max_total_tokens), and scheduled triggers for recurring tasks. Early adopters like OffDeal use hooks for automated logo validation in investment decks. Supported models include Gemini 3.6 Flash (default), 3.5 Flash, and 3.5 Flash-Lite, with explicit model selection via agent_config.
managed agentsenvironment hooksgemini 3.6 flashmax_total_tokensscheduled triggers
5 ways to host the ultimate dinner party with Google Search
Google Search introduces AI-powered features for dinner party planning, leveraging Nano Banana visualization and AI Mode for multi-modal assistance. The system processes trending queries (e.g., 'mahjong dinner party', 'elegant chicken recipes') to provide: (1) tablescape visualization, (2) menu generation with recipe creator metadata, (3) drink pairing recommendations via LLM analysis, (4) app-integrated playlist curation via YouTube Music API, and (5) printable menu design. AI Mode demonstrates contextual understanding by suggesting complementary beverages (e.g., for miso salmon menus) and visual treatments for dish discovery.
nano bananaai modevisual treatmentdrink pairingsmulti-modal
5 ways AI Mode in Search helps you enjoy the real world
Google Search's AI Mode introduces five functionalities to enhance real-world engagement through personalized digital assistance. The system leverages Personal Intelligence for schedule-aware local class recommendations, integrates Canvas for strategy memorization in games, and facilitates event ticket booking via curated options. It also supports outdoor gear shopping with local inventory checks and enables Canva-integrated dinner party invitation design. These features collectively aim to streamline offline activities by providing context-aware, app-integrated assistance, as evidenced by increased user queries for activities like hiking and digital detox.
personal intelligencecanvascanvalocal inventoryschedule-aware
📜 arXiv Papers
No new items today.
📰 Industry Media (2)
Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
The Fly Language Model (FLM) integrates the complete MaleCNS v1.0 fruit fly connectome (166,700 nodes, 25.6M edges) as a fixed reservoir into a frozen LiquidAI LFM2.5-1.2B-Instruct backbone, training only a 278K-parameter readout (0.02% of backbone). Despite a 0.0222 nat/token improvement over the backbone on SmolTalk dialogues (NLL 1.382→1.360), a no-graph control outperformed FLM by 0.000488 nats/token, with bootstrap analysis showing no significant connectome-specific benefit. The reservoir's state decays at 0.6 per token (20-token memory horizon: 3.66e-5), confirming no long-range context contribution. FLM's MIT-licensed implementation runs locally but lacks independent reproducibility.
connectomereservoir computingfrozen backboneperplexitynats/token
Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
The HarnessDev framework evaluates LLMs' ability to engineer their own agent harnesses (execution loops, tools, state) rather than task outputs. Researchers from ByteDance Seed et al. tested six models (Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Max, Seed 2.0 Pro) across 2,207 instances spanning code, search, writing, and ML tasks. Creation phase results showed Opus 4.8 achieved 67.8 avg@3 score (vs human reference 86.2), excelling in writing (84.6 vs 83.7) but lagging in search (52.6 vs 92.2). Evolution phase revealed only 34 of 64 harness modifications generalized to held-out tasks, with executor-specific performance drops (e.g., Opus SWE-Pro score fell from 69.3 to 33.0 under Gemini).
agent harnessself-evalheld-out generalizationexecutor transferdead code
Generated automatically at 2026-09-12 21:13 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.
