Home > DATA > Policy on the Use of AI Tools for CFPS Data

Policy on the Use of AI Tools for CFPS Data

来源:时间:2026-09-02 02:03阅读:

Background

In recent years, the rapid development of large language models (LLM) and online artificial intelligence (AI) tools has exerted a profound impact on the processing and analysis of survey data. Based on the China Family Panel Studies (CFPS) User Agreement and taking into account the potential risks associated with various ways of using AI tools, the Institute of Social Science Survey at Peking University and the CFPS team have formulated the following rules on the use of AI tools.

I. General Provisions

AI tools (including generative-AI models such as Kimi, DeepSeek, Doubao, Zhipu, Qwen, ChatGPT, Claude, Google Gemini, Grok; AI-service platforms accessed via remote API interfaces; browser-based AI assistants or plugins; and any other tools that transmit data to external service providers) shall not be used to process or analyze household-level or individual-level CFPS microdata. This policy applies to all CFPS data users, including users of CFPS public-use data and restricted-use data.

II. Regulatory Basis

Under the current CFPS Data User Agreement, users are not allowed to place CFPS data or processed derived datasets on journal-related websites or third-party platforms. Users are also prohibited from publishing or transferring to third parties all or part of the CFPS data, or its content with modifications, whether under the original name or an alternative title.

Using AI tools to process CFPS data may result in the transmission of data to third-party providers, thereby constituting a violation of the existing CFPS Data User Agreement.

III. Classification and Rules for AI Use

(1) Scope of Prohibited Uses

Modes of AI Tool Usage

Typical Tools and Operations

CFPS Data Requirements

Rationale

1. Direct conversation with chatbots

Directly conversing with AI via web pages or client-side applications, including pasting, typing-in or uploading data files (ChatGPT, Doubao, Kimi, DeepSeek, Claude, Gemini, etc., including their “data analysis/file-upload” functions)

Data users are prohibited from pasting, typing, or uploading household-level or individual-level microdata

Data is sent directly to the service provider’s server and may be retained or used for model training, which amounts to distribution to a third party. Such usage remains prohibited even if the provider states that data will not be retained and will be deleted immediately.

2. Remote API calls

Calling remote AI service interfaces via Python, R, Stata, or other scripts, or through third-party tools (OpenAI API, ERNIE Bot API, Qwen API, etc.)

Data users are prohibited from sending household-level or individual-level microdata as input

Data is transmitted over the network to an external platform, which constitutes third-party data transfer. This prohibition applies regardless of any promise of immediate deletion. Circumventing monitoring through middleware, proxy servers or encryption is also forbidden.

3. Built-in AI assistants and coding Agents within programming platforms/IDEs

Using AI code completion, plugins, or cloud-based coding agents in VS Code, Cursor, RStudio, etc., especially features that can automatically index working directories and read local files

Prohibited from allowing access to or reading of files or directories containing household-level or individual-level microdata

These tools can automatically scan folders and upload local files to external servers, creating a risk of unintentional data leakage. (Usage limited purely to code-writing without access to microdata is permitted; see FAQ 1).

4. Institution-licensed cloud-based AI tools

Cloud services procured under institutional contracts with formal commitments not to retain user data (e.g. enterprisegrade Microsoft 365 Copilot)

Prohibited for use with household-level or individual-level microdata

Data still travels through external servers, with residual risks such as interception during network transmission and cross-reference analysis.

(2) Scope of Permitted Uses

AI-tool usage is allowed for processing the following CFPS-related content, provided that no household-level or individual-level microdata or personally identifiable information is included in any uploaded materials.

1. Publicly available documents accessible without registration, such as technical reports, programming codes and codebooks;

2. Survey metadata (variable labels, value labels);

3. Aggregated population-level statistics that have been deidentified, cannot be used for re-identification of individuals or areas below provincial level.

For any matters not covered above, users are advised to contact the CFPS team (isss.cfps@pku.edu.cn) for consultation. Requests should describe the intended mode of AI-tool use and research objectives. The project team may, as appropriate, refer such cases to the Institute of Social Science Survey, Peking University for collective review. 

IV. Frequently Asked Questions (FAQ)

Question 1: Is it permissible to use LLMs to write analytical code (e.g., Stata, Python, or R scripts) based on CFPS data?

Answer 1: Yes, provided that no household-level or individual-level CFPS microdata is uploaded to the LLMs. Using an LLM to assist in writing analytical code does not involve data transmission and falls within the permitted scope.

Question 2: Is it permissible to share derived analytical outputs (such as model coefficients and summary statistics) with LLMs?

Answer 2: Yes, provided that the shared materials do not contain statistics that could identify households or individuals (e.g., the income of a specific respondent) or areas below provincial level, and that no household-level or individual-level CFPS microdata is uploaded.

Question 3: Is it permissible to use a cloud-hosted coding Agent to automatically index and read the working directory containing CFPS data files?

Answer 3: No. Cloud-hosted coding Agents are capable of autonomously scanning folders and reading the content of local files, which creates a risk of data leakage. Users must disable the automatic-indexing function of such tools, or store CFPS data in a separate directory inaccessible to the Agent. Users should also ensure that this directory is not included in any cloud-sync or cloud-backup services.

Question 4: Is it permissible to process CFPS data by calling LLMs via remote-API services (e.g., OpenAI API, ERNIE Bot API, Qwen API, etc.)?

Answer 4: No. Sending CFPS data to external AI service platforms via remote API services constitutes transmission of data to third parties and violates the data use agreement. Regardless of whether the API provider claims not to retain data, not to reuse data, or to delete data immediately, processing CFPS microdata through internet-connected API interfaces in any manner is prohibited. This provision also applies to batch calls to remote APIs through programming methods (e.g., Python, R, or other scripts), as well as any use of middle layers, proxy services, or technical means to circumvent monitoring of external data transmission. If a locally deployed model is used, it must be ensured that the model and its operating environment do not rely on external cloud services for inference computation or data transmission back to remote servers. If a local model interface silently connects to remote servers in the background, it shall be treated as a remote API service and is strictly prohibited.

Question 5: Is it permissible to process CFPS data using opensource LLMs (such as Llama, Qwen) deployed on a local server?

Answer 5: Yes, subject to strict security requirements. Local deployment must ensure that the model-running environment is physically isolated from the internet or certified by the institutional cybersecurity authority. Data shall be processed exclusively within a closed-loop local environment, with no reliance on external cloud-services for inference or data-transmission. If a local client actually still calls a remote-API, it shall be regarded as a remote API service and is prohibited. Users are recommended to submit their local-deployment plans to the CFPS project team. This provision does not apply to CFPS restricteduse data; no AI tools (including locally-deployed models) may be used with restricted data without prior approval from the project team.

V. New Additions to the CFPS Data User Agreement

The current Chinese and English versions of the Data User Agreement include the following provisions:

Chinese Version:

我承诺不会将CFPS原始数据或其衍生数据集以任何方式(包括上传、导入或API调用)传输至外部人工智能服务平台,包括但不限于生成式AI、数据分析类云端AI工具。本人知悉,使用此类工具涉及向第三方传输数据,将破坏数据安全并违反本协议。 

English Version:

I will not transmit the CFPS original or derived datasets, by any means (including uploading, importing, syncing, or API calls), to any external artificial intelligence service platforms, including but not limited to generative AI tools and cloud-based AI data-analysis tools. I acknowledge that the use of such tools involves transferring data to third parties and thus will compromise data security and constitute a breach of this agreement.


上一篇:

下一篇: