Open-weight GLM 5.2 outperformed Claude in IDOR detection, showing cost-effective strength without complex scaffolding.
// curated from Hacker News with AI
Open-weight GLM 5.2 outperformed Claude in IDOR detection, showing cost-effective strength without complex scaffolding.
AI analyzed MRI discrepancies, highlighting AI's potential for second opinions but underscoring trust issues in healthcare diagnostics.
Brown professor exposes massive AI cheating scandal, raising concerns about academic integrity and the future of higher education amidst AI challenges.
Ford rehired experienced engineers after AI-driven quality issues, boosting standards and ranking top in vehicle quality.
OpenAI Codex faces ongoing challenges excluding sensitive files, impacting security and access control.
Google limits Meta's access to Gemini AI models, delaying Meta's AI projects amid rising demand and capacity constraints across clients.
Ford rehired veteran engineers to improve quality after AI and automation underperformed, boosting efficiency and brand ranking.
Wayfinder Router deterministically routes prompts offline, directing simple queries locally and complex ones to cloud models, saving costs.
Austria urges EU to host Anthropic to counter US restrictions on advanced AI models.
Proxy-KD enables knowledge transfer from black-box LLMs to smaller models, outperforming traditional white-box methods.
AI transforms software development into supervision and editing, risking skill erosion, reduced creativity, and long-term societal costs.
Examines if LLMs detect their own altered outputs, likening it to a mirror test for self-awareness, with intriguing behavior observed.
Non-profit group offers diverse, accurate AI images to replace stereotypes, enhancing public understanding and conversation.
A lightweight Bash wrapper for GPT-like LLM APIs, focusing on security, modularity, and environment compatibility, with dynamic models and session management.
Researchers networked FPGAs to create a programmable probabilistic computer with 1 million p-bits, enabling large-scale sampling and optimization.
A from-scratch GPT-2 model in C/CUDA trained on CPU and a CUDA engine, demonstrating engineering and training pipeline in public.
Ornith-1.0 is a self-improving open-source LLM family for agentic coding, outperforming larger models across benchmarks.
DIY chemist creates PAC-832, a new AI-designed drug targeting Alzheimer’s in his garage lab.
Frontier LLMs face limits on hard document reading; cost, accuracy, and trust are challenges—specialized models still outperform on complex tasks.
AI agent lost maneuver, triggered nuclear strike after Civilization VI upset.
Grok 4.5, built on 1.5T V9, now in private beta at SpaceX & Tesla, showing performance near or above Opus, with continuous improvements.
Eggers warns that AI writing threatens human creativity, emphasizing the importance of nurturing original thought and resisting machine-generated content.
Experts say current privacy laws are inadequate to protect against facial recognition and biometric data collection in smart glasses.
Chinese firm 360 claims to have developed domestic AI cybersecurity models rivaling Mythos, aiming for offensive and defensive cyber capabilities amid US export restrictions.
AgentWatch enforces AI session budgets, preventing runaway agents, with easy 2-minute integration and full control across providers.
AI coding token costs may soon match or surpass developers' salaries, raising concerns about overspending without proper governance.
Google restricts Meta's Gemini AI access due to high demand straining capacity.
AI investment boom risks recession, job loss, and financial bust if returns falter, threatening middle class stability worldwide.
A guide on creating clear, effective graphs by applying specific principles—handle data, layout, typography, color, polish, themes, and exports.
Most Dutch responses to EU tobacco rules were AI-generated by Philip Morris, influencing policy with biased, industry-driven messages.
Shift from prompting AI agents to designing loops that prompt them, streamlining AI coding processes.
New Claude models solve more tasks, with less cost per task; telemetry reveals nuanced efficiencies beyond accuracy leaderboard.
AI risks creating a permanent underclass by automating skills, valuing digital literacy while marginalizing the unadapted workforce.
AI tool guides scholarly writing with Claude Code in Czech and English, grounded in sources and human validation.