Studies in Science of Science ›› 2026, Vol. 44 ›› Issue (8): 1603-1612.

Previous Articles     Next Articles

Risks Intellectual Property Risks and Compliance Strategies on the User Side of Open-Source Foundation Models: A Legal Perspective on License Breach under Open Source Agreements

  

  • Received:2025-05-26 Revised:2025-09-24 Online:2026-08-15 Published:2026-08-15

开源大模型使用端的知识产权风险及合规对策

韩硕1,丁天曲2   

  1. 1. 湘潭大学法学学部知识产权学院
    2. 湘潭大学
  • 通讯作者: 韩硕
  • 基金资助:
    开放创新范式下知识共享机制研究

Abstract: Current research on the construction and governance of open-source large model ecosystems primarily focuses on macro-level top-level design, with a notable lack of micro-level risk identification and compliance strategies. In the open-source AI ecosystem, diverse actors increasingly participate in model deployment, creating an urgent need for theoretical guidance to support industry compliance practices.This study systematically reviews representative open-source licenses in the global large model domain, identifies intellectual property risks potentially triggered by license breaches, and employs an interdisciplinary "Law-Technology-Intelligence" analytical framework to propose countermeasures.The primary causes of license breach risks include insufficient clearance of rights-related information during the data preprocessing stage, lack of typological understanding of open-source license provisions during application management, misaligned risk recognition, and inadequate control over derivative outputs.At the data preprocessing (upstream) stage, enterprises should implement a full-cycle rights information clearance mechanism integrating technical filtering, responsibility clarification, and data compliance.At the model management (midstream) stage, organizations must account for the typological traits and risk profiles of permissive licenses,restrictive-permissive licenses, and copyleft licenses, and accordingly establish differentiated compliance governance systems.At the content output (downstream) stage, compliance management should encompass four dimensions: data handling, model training, generation control, and user supervision.

摘要: 大模型开源生态建设及分享治理的研究集中于顶层设计,缺乏微观层面风险的识别及合规对策研究,开源AI生态下,各类主体泛化参与模型部署,亟需为产业合规实践提供理论指引。梳理世界范围大模型领域的典型开源协议,识别因协议违约可能导致的知识产权风险,运用“法律-技术-情报”交叉融合范式分析解决对策。数据预处理阶段的权利信息清理不足、应用管理阶段的开源协议条款的类型化认知不清、风险识别错位以及输出阶段衍生内容的控制缺失是开源协议违约风险的主要致因。前端训练数据处理层面,企业应建立“技术过滤+权责明示+数据合规”的全周期权利信息清理机制。中端智能技术管理层面,需注重宽松开源协议、限制性宽松开源协议与感染性开源协议的类型学特征和风险特点,分别设置合规管理体系。后端输出内容阶段,需从数据处理、模型训练、生成控制与用户监督四个维度展开合规管理。