Learning desk

MultiAgent EDU StackGather good sources. Teach what matters.
T5TauricResearch/TradingAgentsT5A Man Who Invented Modern AI (Before Everyone Else) – Jürgen Schmidhuber [video]T5GPT-4 finished training four years ago todayT5AI Settles a 25 Year-Old Problem We Left BehindT5What it was like working on LLMs and security at Meta (2022-2026)T5Ask HN: How do you go from writing code to deploying with agents?T5What Happened: OpenAI and HuggingFaceT5Apple says Mac users in China can connect to Alibaba's Qwen AI serviceT5Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on SonnetT5The AI Apocalypse Is HereT3Auto mode is now the default in Claude Code for Pro, Max, and Team plansT5Show HN: Tura – Build agent that uses 80% less token and delivers better resultsT5TauricResearch/TradingAgentsT5A Man Who Invented Modern AI (Before Everyone Else) – Jürgen Schmidhuber [video]T5GPT-4 finished training four years ago todayT5AI Settles a 25 Year-Old Problem We Left BehindT5What it was like working on LLMs and security at Meta (2022-2026)T5Ask HN: How do you go from writing code to deploying with agents?T5What Happened: OpenAI and HuggingFaceT5Apple says Mac users in China can connect to Alibaba's Qwen AI serviceT5Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on SonnetT5The AI Apocalypse Is HereT3Auto mode is now the default in Claude Code for Pro, Max, and Team plansT5Show HN: Tura – Build agent that uses 80% less token and delivers better results
← Dispatches

Shifting the Unit of Safety: From Model to System in the Generative and Agentic Era

Primary research

#1490

T1new
Topic
unassigned (set during synthesis)
First seen
2026-08-09 07:16:03
Last seen
2026-08-09 07:16:03

Source raw items (1)

  • Semantic Scholar2026-08-09 07:15:26
    Shifting the Unit of Safety: From Model to System in the Generative and Agentic Era

    For a decade, responsible AI at internet scale rested on a reassuring assumption: that risk lives primarily in a model, it is a unit you can isolate, and that privacy, fairness and safety is therefore something you certify at model level, before launch. At LinkedIn, where AI shapes access to jobs and economic opportunity for over a billion members, we built our fairness, privacy, and explainability assurance on exactly this foundation, and it worked. Then generative AI dissolved the unit, and agentic AI dissolved the boundary. When a single product's AI harness composes prompts, tools, retrieved context, and multi-turn reasoning, harm no longer announces itself simply at the model layer or at deployment time. It emerges across countless dynamic call patterns at the system level, in context of the product usage, precisely where a pre-launch gate for models cannot see it. This talk is about the architectural and philosophical shift this forces, and how a Responsible AI & Governance organization can navigate the resulting tension: the pressure to ship AI at the speed of competition, against a risk surface that evolves with an expanding suite of AI harness - prompts, tools, memory and orchestration. I will argue that safety must move from a one-time high fidelity gate to a continuous risk management property, and share how we operationalized that conviction. I will cover in-context red-teaming, which stress-tests a model inside the specific task it performs rather than in the abstract; and our move toward agentic safety, where we treat privacy as contextual integrity, enforce fine-grained access control, and push safety checks down to every ingress and egress of a core AI system, not just its final answer. Throughout, an emerging regulatory landscape (the EU AI Act and its post-market monitoring obligations) serves not as a constraint to satisfy but as a design pattern to internalize. My hope is to leave this audience at the intersection of industry-scale challenges and academic rigor with a concrete reframing and a set of open problems: How do we certify a system whose call patterns itself are highly dynamic? What does a meaningful safety guarantee mean for an agent that plans? And how do we measure all of this without collecting the very sensitive data we are trying to protect? These are LinkedIn's daily engineering reality; they are also, I will argue, some of the most consequential open research questions in the field of scalable AI Safety.