Daily Digest — 2026-07-20
4 items · 4 industry media
🏛️ Research Labs
No new items today.
📜 arXiv Papers
No new items today.
📰 Industry Media (4)
Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep
Perplexity AI introduces WANDR, an open benchmark for evaluating research agents on wide-and-deep evidence collection tasks. The benchmark comprises 500 realistic tasks structured as qualification key hierarchies, requiring agents to discover entities and validate claims with cited evidence. Tasks are generated via a semi-automated pipeline, with grading based on re-fetched URL verification. Evaluation of six systems shows Perplexity's Search as Code leading (0.363 soft F1), though all systems struggle with discovery and complete evidence extraction, highlighting structural bottlenecks in scalable knowledge work.
research agentsevidence collectionqualification key hierarchyreference-free gradingsearch as code
10 Open-Source No-Code AI Platforms for Building LLM Apps, RAG Systems, and AI Agents
The article surveys 10 open-source no-code/low-code platforms for developing LLM applications, RAG systems, and AI agents, comparing their interfaces, licensing, and use cases. Platforms include HKUDS AutoAgent (MIT) for zero-code agent creation via natural language, Mintplex Labs AnythingLLM (MIT) for private RAG systems, and InfiniFlow RAGFlow (Apache-2.0) for document parsing. Key findings highlight mature visual/no-code tooling for retrieval and agent workflows, with varying license restrictions (MIT, Apache-2.0, or proprietary modifications) and specialization in document accuracy, production monitoring, or automation integration.
rag systemsmulti-agent workflowsdocument parsingllm orchestrationvisual workflow builder
Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost
The article compares three open-weight trillion-scale Mixture-of-Experts (MoE) models—Kimi K3 (2.8T), DeepSeek V4 Pro (1.6T), and GLM-5.2 (744B)—on capability, licensing, and serving costs. All models feature 1M-token context windows and target long-horizon coding tasks. Benchmark results from Artificial Analysis Intelligence Index show Kimi K3 leads (score 57), followed by GLM-5.2 (51) and DeepSeek V4 Pro (44). DeepSeek V4 Pro offers the lowest serving cost ($0.04/task), while GLM-5.2 provides the fastest inference (168 tokens/sec). Licensing varies, with DeepSeek and GLM offering MIT-licensed weights, while Kimi K3 remains API-only until July 2026.
mixture-of-expertscontext windowbenchmarkserving costopen-weight
Fine-Tuning Qwen3 with LoRA Using NVIDIA NeMo AutoModel: A Complete Single-GPU Google Colab Workflow Tutorial
The tutorial presents a workflow for parameter-efficient fine-tuning of Qwen3-0.6B using LoRA via NVIDIA NeMo AutoModel on a single-GPU Google Colab instance. It demonstrates environment setup, YAML recipe adaptation for constrained resources (reducing batch size to 4, global batch to 8, and limiting to 40 steps), and evaluation comparing base and fine-tuned models on HellaSwag. Results show successful LoRA checkpoint generation and integration through PEFT, while maintaining compatibility with distributed scaling via NeMo's configuration-driven architecture.
loranemo automodelqwen3parameter-efficient fine-tuninghellaswag
Generated automatically at 2026-07-19 20:07 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.
