CN117560159A - A secure and fair federated learning privacy-preserving aggregation system and method - Google Patents

A secure and fair federated learning privacy-preserving aggregation system and method Download PDF

Info

Publication number
CN117560159A
CN117560159A CN202311527204.6A CN202311527204A CN117560159A CN 117560159 A CN117560159 A CN 117560159A CN 202311527204 A CN202311527204 A CN 202311527204A CN 117560159 A CN117560159 A CN 117560159A
Authority
CN
China
Prior art keywords
model
message
local
models
terminal equipment
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
CN202311527204.6A
Other languages
Chinese (zh)
Inventor
张文芳
蒲庆
张倚涛
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Southwest Jiaotong University
Original Assignee
Southwest Jiaotong University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Southwest Jiaotong University filed Critical Southwest Jiaotong University
Priority to CN202311527204.6A priority Critical patent/CN117560159A/en
Publication of CN117560159A publication Critical patent/CN117560159A/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F21/00Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
    • G06F21/60Protecting data
    • G06F21/62Protecting access to data via a platform, e.g. using keys or access control rules
    • G06F21/6218Protecting access to data via a platform, e.g. using keys or access control rules to a system of files or objects, e.g. local or distributed file system or database
    • G06F21/6245Protecting personal data, e.g. for financial or medical purposes
    • G06F21/6254Protecting personal data, e.g. for financial or medical purposes by anonymising data, e.g. decorrelating personal data from the owner's identification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/213Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods
    • G06F18/2135Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods based on approximation criteria, e.g. principal component analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/23Clustering techniques
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L63/00Network architectures or network communication protocols for network security
    • H04L63/04Network architectures or network communication protocols for network security for providing a confidential data exchange among entities communicating through data packet networks
    • H04L63/0407Network architectures or network communication protocols for network security for providing a confidential data exchange among entities communicating through data packet networks wherein the identity of one or more communicating identities is hidden
    • H04L63/0421Anonymous communication, i.e. the party's identifiers are hidden from the other party or parties, e.g. using an anonymizer
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L67/00Network arrangements or protocols for supporting network services or applications
    • H04L67/01Protocols
    • H04L67/10Protocols in which an application is distributed across nodes in the network
    • H04L67/104Peer-to-peer [P2P] networks
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L9/00Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols
    • H04L9/002Countermeasures against attacks on cryptographic mechanisms
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L9/00Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols
    • H04L9/32Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols including means for verifying the identity or authority of a user of the system or for message authentication, e.g. authorization, entity authentication, data integrity or data verification, non-repudiation, key authentication or verification of credentials
    • H04L9/3247Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols including means for verifying the identity or authority of a user of the system or for message authentication, e.g. authorization, entity authentication, data integrity or data verification, non-repudiation, key authentication or verification of credentials involving digital signatures
    • H04L9/3255Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols including means for verifying the identity or authority of a user of the system or for message authentication, e.g. authorization, entity authentication, data integrity or data verification, non-repudiation, key authentication or verification of credentials involving digital signatures using group based signatures, e.g. ring or threshold signatures
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L9/00Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols
    • H04L9/40Network security protocols

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Security & Cryptography (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • General Engineering & Computer Science (AREA)
  • Signal Processing (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Software Systems (AREA)
  • Computer Hardware Design (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • Medical Informatics (AREA)
  • Computing Systems (AREA)
  • General Health & Medical Sciences (AREA)
  • Bioethics (AREA)
  • Databases & Information Systems (AREA)
  • Mathematical Physics (AREA)
  • Computer And Data Communications (AREA)

Abstract

The invention discloses a safe and fair federal learning privacy protection aggregation system and method, relates to the technical field of federal learning, and solves the problems of low calculation efficiency and unfairness to a few clients in the existing distributed cross-equipment federal learning scheme; the cloud aggregation server is used for identifying and removing malicious models, and updating the global model after aggregating the benign models; the edge computing node is used for realizing anonymization of the model at the communication level; the terminal equipment of the Internet of things is used for carrying out model training on a local data set to obtain local model update and then encrypting and uploading the local model update to a corresponding edge computing node; n training clients develop distributed cross-equipment transverse federation learning with the assistance of M edge computing nodes; the method has higher malicious model detection rate, and maintains aggregation fairness for a few group models, so that the global model can learn knowledge from diversified data, and the quality of the global model is improved.

Description

Safe and fair federal learning privacy protection aggregation system and method
Technical Field
The invention relates to the technical field of federal learning, in particular to a safe and fair federal learning privacy protection aggregation system and method.
Background
In recent years, based on the unique privacy protection attribute of federal learning, introducing federal learning into the internet of things has become an effective way to solve the data security problem of the internet of things. Existing studies indicate that federally learned distributed architecture is vulnerable to privacy and security attacks. More specifically, gradient updates uploaded by edge clients are likely to be subject to inference attacks by malicious analysts or servers, thereby obtaining private information of the user's local data. More obvious is security attack, especially in the distributed network application scenario of B2C, compared with the aggregation server, the edge client device is more easily attacked by malicious and controlled by adversary, and then the global model is blocked from converging (bayer attack) or generating wrong models (poisoning attack) by submitting the malicious model update mode.
While for the defense of model poisoning attack, three strategies exist at present: 1. outliers are removed from the local model set based on distance, similarity between models. 2. The aggregation server tests the received model by its own set of tests. 3. The reputation system of the client is built by means of a blockchain, and the reputation of the client is dynamically adjusted based on the quality of the model uploaded by the client each time. However, there are certain problems with all three defense strategies described above. Based on reputation evaluation, the aim of restraining benign clients can be achieved, but after a malicious attacker invades a certain client, model poisoning attack can be directly developed to influence aggregation of the global model. Based on the server-side model quality detection scheme, the aggregation server is required to have a corresponding test data set, which is unreasonable in many application scenarios. But the anomaly detection means based on the model distance only has poor performance effect on the adaptive attack developed by the adversary with priori knowledge.
On the other hand, how to guarantee fairness of the anomaly model detection algorithm is also a problem. In a distributed scene of the internet of things, data collected by part of terminal equipment may show the characteristics of non-independent and same distribution, and obvious and reasonable differences exist between a local model submitted by the part of equipment and a model submitted by most of equipment. However, existing anomaly detection algorithms tend to ignore this reasonable variance, remove all outliers, which is unfair to the model of a minority of benign populations, nor meets the original intent of using differentiated data to promote global model quality. It is therefore necessary to identify and preserve differentiated good models while detecting malicious models.
From the above analysis, the main problems with existing distributed cross-device (federal learning schemes) are: 1. it is difficult to meet both the requirements of model privacy and aggregate security. Even if schemes based in part on homomorphic encryption and third party server-assisted detection are satisfied, their communication protocols tend to be complex and computationally inefficient, and secure assumptions about the servers are strong. 2. Fairness detection of the anomaly model is not achieved. The anomaly detection scheme based on Euclidean distance and cosine similarity often fails in high latitude, ignores the diversity of models, and is unfair to a few clients.
Disclosure of Invention
In order to solve the problems in the prior art, the invention aims to provide a safe and fair federal learning privacy protection aggregation system and method, and aims to solve the problems of low calculation efficiency and unfairness to a few clients in the existing distributed cross-equipment federal learning scheme.
A safe and fair federal learning privacy protection aggregation system comprises a cloud aggregation server, edge computing nodes, an Internet of things terminal device and N training clients for developing distributed cross-device transverse federal learning under the assistance of M edge computing nodes; wherein,
the cloud aggregation server is used for carrying out fair anomaly detection on the collected local models, identifying and removing malicious models, updating the global model after aggregating the benign models, and then sending the global model to the intelligent terminal equipment.
The edge computing node is used for packaging the collected local model and transmitting the packaged local model to the anonymous communication system, so that anonymization of the model is realized at the communication layer.
The terminal equipment of the Internet of things is used for collecting data in a local area network, then model training is carried out on a local data set to obtain local model update, anonymity of the model is achieved on a message layer by utilizing a ring signature technology, and finally the local model update is uploaded to a corresponding edge computing node.
A secure and fair federal learning privacy preserving aggregation method comprising:
s1, initializing: terminal equipment C i Apply for Ring signing Key pair { sk i ,pk i Negotiating public key encryption algorithm g, hash function H and symmetric encryption algorithm E, edge node E j Apply for signing Key Pair { SK j ,PK j Establishing a P2P network among M edge nodes, and publishing a key pair { SK by a cloud server M ,PK M -and send the initialisation Model to all terminal devices;
s2, a local training stage: terminal equipment C i Receiving the initialized Model and performing training locally to obtain local Model update U i Updating the local model U i Encryption and attachment of a ring signature sigma i As model message M i Transmitting to an edge node;
s3, an anonymous communication system transmission stage: edge node E j Generating a message set G after receiving a certain number of model messages j ={M i |i∈S j S, where S j For E j The terminal equipment set under management forwards the message set to another edge node or cloud server according to the anonymous communication protocol;
s4, malicious model detection: model message M in received message set by cloud server i Performing ring signature verification, performing data dimension reduction on the model message, and performing homogeneous clustering technologyDetecting a malicious model;
s5, a global model updating stage: and (3) after all the malicious models detected in the step S4 are removed, the remaining benign models are aggregated, and global model updating is completed.
Preferably, the S2 includes:
s201, terminal equipment C i After the local model is trained, the model local update U is obtained i Simultaneously generating a random number N i According to public key PK of cloud server M Encryption to obtain ciphertext PK M (U i ||N i )。
S202, terminal equipment C i Randomly picking a part of the participants' signature public key as a public key set R and calculating a ring signature
σ i =Sign ring (H(U i ||N i ) And send M i =(PK M (U i ||N i ),σ i ) As model messages to the corresponding edge computing node E j }。
Preferably, the anonymous communication protocol in S3 includes:
s301, edge node E α Collecting a plurality of M i Combined into a message set G α ={M i |i∈S α Edge node E α Will G α With (G) α ,SK α (G α ||E β ) In the form of a) to another edge node E β
S302、E β Making a random selection, i.e. submitting to the cloud server with a probability of 1-p (G α ,SK γ (G α Server), and forwards the message with probability of p (G) α ,SK β (G α ||E γ ) For another edge node E) γ
S303、E γ A decision will be made with the same probability whether to submit the message (G α ,SK γ (G α Server) to the cloud Server, or continue forwarding messages;
s304, the cloud server receives the message (G) α ,SK j (G α Server), and thenFor message set G α ={M i |i∈S α Each model message M in } i =(PK M (U i ||N i ),σ i ) Performing decryption calculation and verifying the ring signature sigma i To ensure the correctness of each model message M i Indeed by legal terminal equipment C i The generation; after the verification of the ring signature is passed, the random number N is checked i Ensure M i Not repeated messages.
Preferably, the S4 includes:
s401, the cloud server first receives all the received model messages M i =(PK M (U i ||N i ),σ i ) And performing ring signature verification to exclude all illegal model messages.
S402, constructing a legal model update set { U } i I is more than or equal to 1 and less than or equal to N, and the principal component analysis is performed to realize the dimension reduction of the data. And carrying out micro aggregation on the data subjected to dimension reduction to realize homogeneous clustering. And finally, carrying out malicious model detection in each homogeneous class.
The beneficial effects of the invention include:
an anonymous submission mechanism of a model is designed based on a ring signature technology and a P2P network, so that model update has non-connectivity, and sensitive information of local data is prevented from being stolen by a cloud server to develop model reasoning attack. And then, aiming at model updating in a plaintext state, realizing abnormal model fairness detection by utilizing principal component analysis and a micro-aggregation algorithm. The security analysis shows that the proposed scheme can truly realize anonymity submission of the model, and the privacy of local data is ensured through unlinkability. Experimental evaluation proves that the detection scheme provided by the invention has higher malicious model detection rate, and has aggregation fairness for few models, so that the global model can learn knowledge from diversified data, and the quality of the global model is improved.
Drawings
Fig. 1 is a diagram of a secure and fair federal learning privacy protection aggregation system architecture according to embodiment 1.
Fig. 2 is a flowchart of a secure and fair federal learning privacy preserving aggregation method according to embodiment 1.
Fig. 3 is a diagram of the present scheme related to embodiment 1 in comparison to the overhead of the prior art in terms of computation and communication.
Fig. 4 is a comparison of the model accuracy of the scheme according to example 1 with the prior art at a malicious model ratio of 10%.
Fig. 5 is a comparison of the model accuracy of the scheme according to example 1 with the prior art at a malicious model ratio of 20%.
Detailed Description
For the purposes of making the objects, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application, and it is apparent that the described embodiments are only some embodiments of the present application, but not all embodiments. Thus, the following detailed description of the embodiments of the present application, as provided in the accompanying drawings, is not intended to limit the scope of the application, as claimed, but is merely representative of selected embodiments of the application. All other embodiments, which can be made by those skilled in the art based on the embodiments of the present application without making any inventive effort, are intended to be within the scope of the present application.
Example 1
Specific embodiments of the present invention will be described in detail below with reference to fig. 1-5;
a safe and fair federal learning privacy protection aggregation system is shown in figure 1, and comprises a cloud aggregation server, edge computing nodes, an Internet of things terminal device and N training clients for developing distributed cross-device transverse federal learning under the assistance of M edge computing nodes; wherein,
the cloud aggregation server is used for carrying out fair anomaly detection on the collected local models, identifying and removing malicious models, updating the global model after aggregating the benign models, and then sending the global model to the intelligent terminal equipment.
The edge computing node is used for packaging the collected local model and transmitting the packaged local model to the anonymous communication system, so that anonymization of the model is realized at the communication layer.
The terminal equipment of the Internet of things is used for collecting data in a local area network, then model training is carried out on a local data set to obtain local model update, anonymity of the model is achieved on a message layer by utilizing a ring signature technology, and finally the local model update is uploaded to a corresponding edge computing node.
When the method is described, the cloud aggregation server is briefly described as a cloud server, the edge computing nodes are briefly described as edge nodes, and the terminal equipment of the Internet of things is briefly described as terminal equipment; for ease of understanding, the main symbol and variable list to which this embodiment relates is presented:
TABLE 1 variable sign and definition
A secure and fair federal learning privacy preserving aggregation method comprising:
s1, initializing: terminal equipment C i Apply for Ring signing Key pair { sk i ,pk i Negotiating public key encryption algorithm g, hash function H and symmetric encryption algorithm E, edge node E j Apply for signing Key Pair { SK j ,PK j Establishing a P2P network among M edge nodes as a bottom communication architecture of an anonymous communication protocol; cloud server published key pair { SK M ,PK M -and send the initialisation Model to all terminal devices;
s2, a local training stage: terminal equipment C i Receiving the initialized Model and performing training locally to obtain local Model update U i Updating the local model U i Encryption and attachment of a ring signature sigma i As model message M i Transmitting to an edge node;
s201, terminal equipment C i After the local model is trained, the model local update U is obtained i Simultaneously generating a random number N i According to public key PK of cloud server M Encryption to obtain ciphertext PK M (U i ||N i )。
S202, terminal equipment C i Randomly picking a part of the participants' signature public key as a public key set R and calculating a ring signature
σ i =Sign ring (H(U i ||N i ) And send M i =(PK M (U i ||N i ),σ i ) As model messages to the corresponding edge computing node E j The method comprises the steps of carrying out a first treatment on the surface of the Specific:
each intelligent terminal equipment C i Freely selecting a set R, C containing R, (r.ltoreq.N) ring signatures from the public key sets of all intelligent devices i Public key pk of a signature of (a) i Is the s-th public key in the set R, and has s which is more than or equal to 1 and less than or equal to R. Calculating a symmetric key k=h (U i ||N i ) Then at {0,1 }) b Inner random selection of r-1 x j (j. Noteq. S) as one-way trapdoor function input to other public key counterpart devices in set R, i.e., y j =g j (x j ) Select {0,1} b Random number v in the range is taken as an initial value, and a combination function is constructed:by solving the above combined function, y can be obtained s . Finally C i Calculating the input of the own one-way trapdoor function>Then output H (U) i ||N i ) Ring signature sigma of (2) i ={pk 1 ,pk 2 ,…pk r ,v,x 1 ,x 2 ,…x r }. Ring signature sigma i In the identification information H (U i ||N i ) At the same time of legal validity, the true signer C i The identity information of the set R is hidden in a plurality of intelligent devices, so that the anonymity of the message is realized. Finally C i Message M i =(PK M (U i ||N i ),σ i ) Uploading to the corresponding edge computing node E j
S3, an anonymous communication system transmission stage: edge node E j Generating a message set G after receiving a certain number of model messages j ={M i |i∈S j S, where S j For E j The terminal equipment set under management forwards the message set to another edge node or cloud server according to the anonymous communication protocol; wherein the anonymous communication protocol includes:
s301, edge node E α Collecting a plurality of M i Combined into a message set G α ={M i |i∈S α Edge node E α Will G α With (G) α ,SK α (G α ||E β ) In the form of a) to another edge node E β
S302、E β Making a random selection, i.e. submitting to the cloud server with a probability of 1-p (G α ,SK γ (G α Server), and forwards the message with probability of p (G) α ,SK β (G α ||E γ ) For another edge node E) γ
S303、E γ A decision will be made with the same probability whether to submit the message (G α ,SK γ (G α Server) to the cloud Server, or continue forwarding messages;
s304, the cloud server receives the message (G) α ,SK j (G α Server), then for message set G α ={M i |i∈S α Each model message M in } i =(PK M (U i ||N i ),σ i ) Performing decryption calculation and verifying the ring signature sigma i To ensure the correctness of each model message M i Indeed by legal terminal equipment C i The generation; after the verification of the ring signature is passed, the random number N is checked i Ensure M i Not repeated messages.
S4, malicious model detection: model message M in received message set by cloud server i Performing ring signature verification, performing data dimension reduction on the model message, and performing malicious model detection through a homogeneous clustering technology;
s401, after waiting for a set time, the server will eliminate all the collected messagesRest M i =(PK M (U i ||N i ),σ i ) And (5) performing ring signature verification. According to sigma i Calculating all y recorded in R j =g j (x j ) And symmetric key k=h (U i ||N i ) Then verify the equationWhether or not it is. If true, mark M i Is from intelligent equipment C of the Internet of things i Is a legal model message of (1).
S402, after ring signature verification of all model messages is completed, the cloud server needs to verify { U } i I is more than or equal to 1 and less than or equal to N is used as principal component analysis, so that dimension reduction of model data is realized. Server construction matrixWherein U is i =[u 1i ,u 2i ,...,u mi ]. Firstly, carrying out zero mean value normalization processing on model data, and calculating a model mean value:
calculating model data variance:
performing data normalization calculations:
the data is preprocessed to obtain a matrix G', and then a covariance matrix of the model data is calculated:
the covariance matrix X (G) needs to be diagonalized next to the form:
wherein lambda is m Is the eigenvalue of covariance matrix, and the corresponding eigenvector is e m ∈R m×1 And has lambda 1 >λ 2 >L>λ m . Finally, the first d feature vectors { e } can be selected i I 1 is less than or equal to i is less than or equal to d, and the original m-dimension vector is reduced to d-dimension.
Wherein Y is a matrix of d rows and N columns. The model vector after dimension reduction can well represent the main characteristics of the original model update and can be used for subsequent micro-clustering.
S403, performing micro-aggregation classification on all column vectors (namely principal component vectors updated by each model) in the matrix Y. In this embodiment, the clustering process is completed by using a heuristic fixed-length classification algorithm MDAV, and finally, a plurality of homogeneous classes with k record sizes are obtained. Then dividing the original model update set { U } according to the homogeneity of the main components of the model i I1 is not less than i is not less than N, and is divided into a plurality of micro clusters. In this way, the maximum similarity of models in classes is achieved, and the maximum difference of models between classes is achieved. Because the models of the majority population and minority population are clustered separately, the anomaly models in the homogeneous class are more likely to be malicious models issued by an attacker, since even for the minority population, these anomaly updates are not common. And finally, eliminating abnormal models based on the similarity among the models in each micro-cluster, and completing fairness detection of the malicious models.
S5, a global model updating stage: and (3) after all the malicious models detected in the step S4 are removed, the remaining benign models are aggregated, and global model updating is completed. After eliminating malicious updates in all homogeneous classes, the server will aggregate the benign models remaining in all clusters, and this embodiment uses the federal averaging algorithm to complete the global model update. The updated global model is sent to the intelligent equipment of the Internet of things to carry out subsequent training.
According to the technical scheme of the embodiment, the experimental evaluation comprises the following steps:
the MNIST dataset was used as a training dataset, which was a picture set of 0-9 handwritten digits, each digit having approximately 6000 pictures, the entire training set having 6 ten thousand pictures, and the test set having 2 ten thousand pictures. In the experiment, N=30 intelligent edge devices are selected, M=6 edge computing nodes participate in training, in order to simulate the updating of a few-group differentiated model, the digital 0 picture is only divided into 3 intelligent edge devices (digital 1 and digital 2 are the same), and the rest 21 devices halve the scrambled digital value of the digital value of 3-9, so that the similarity experiment between the multiple-group models is ensured to set the batch size=32, the global training is carried out for 40 rounds, the local training is carried out for 3 rounds, and in addition, in order to accelerate the training task, the malicious poisoning attack and anomaly detection are only carried out on a linear layer which finally contains 330 parameters.
First, the overhead of the present embodiment in terms of computation and communication is analyzed: as shown in fig. 3, as P increases, the number of times the model set is transmitted in the P2P network increases, and the better the anonymity effect of the model set, the stronger the privacy of the model in the model set, but the communication overhead and the calculation overhead also increase. Compared with the Paillier scheme [1] encryption model, when the value p=0.5, the communication cost is almost consistent, the calculation cost is smaller, and the model privacy protection method of the embodiment scheme can enable the subsequent malicious model detection to be more accurate.
Secondly, analyzing a malicious model detection mechanism: when the malicious model proportions are set to 10% and 20%, the effects of the embodiment scheme, DNC scheme [2], bulyan scheme [3], krum [4] scheme and Emd [5] scheme are compared. The DNC scheme is an anomaly detection scheme based on matrix spectrum analysis, the buman, krum scheme is a bayer fault-tolerant aggregation scheme, and the Emd scheme uses EMD distances that are more resolved for high-latitude models for malicious model identification. In order to adapt to the gradient update value which is gradually reduced under the condition of normal convergence of the model, the malicious attack applied by the embodiment can gradually increase the attack intensity along with the advancement of global training.
As shown in fig. 4 and 5, even when the malicious model proportion is 10%, the accuracy of the baseline scheme and the DNC scheme model is poor, and the global model does not converge due to the malicious poisoning attack. Meanwhile, it can be observed that the global models trained by the Krum scheme, the Bulyan scheme, the Emd scheme and the embodiment scheme finally tend to converge, and a better model effect is obtained. And comparing it can be seen that when the proportion of malicious models is larger, the benign models for aggregation are reduced, so that model accuracy of all convergence schemes is reduced under the same number of rounds.
The specific comparison scheme is as follows:
[1]Aono Y,Hayashi T,Wang L,et al.Privacy-preserving deep learning via additively homomorphic encryption[J].IEEE Transactions on Information Forensics and Security,2017,13(5):1333-1345.
[2]Shejwalkar V,Houmansadr A.Manipulating the byzantine:optimizing model poisoning attacks and defenses for federated learning[C].28th Annual Network and Distributed System Security Symposium(NDSS),ELECTR NETWORK:NDSS,2021.
[3]Guerraoui R,Rouault S.The hidden vulnerability of distributed learning in byzantium[C].35th International Conference on Machine Learning(ICML),Stockholm,Sweden:ACM,2018:3521-3530.
[4]Fang M,Cao X,Jia J,et al.Local model poisoning attacks to byzantine-robust federated learning[C].Proceedings of the 29th USENIX Conference on Security Symposium:USENIX Association,2020:1623-1640.
[5]Wang J,Xu G,Lei W,et al.CPFL:an effective secure cognitive personalized federated learning mechanism for industry 4.0[J].IEEE Transactions on Industrial Informatics,2022,18(10):7186-7195.
the foregoing examples merely represent specific embodiments of the present application, which are described in more detail and are not to be construed as limiting the scope of the present application. It should be noted that, for those skilled in the art, several variations and modifications can be made without departing from the technical solution of the present application, which fall within the protection scope of the present application.

Claims (5)

1. The safe and fair federal learning privacy protection aggregation system is characterized by comprising a cloud aggregation server, edge computing nodes, internet of things terminal equipment and N training clients for developing distributed cross-equipment transverse federal learning under the assistance of M edge computing nodes; wherein,
the cloud aggregation server is used for carrying out fair anomaly detection on the collected local models, identifying and removing malicious models, updating the global model after aggregating the benign models, and then sending the global model to the intelligent terminal equipment;
the edge computing node is used for packaging the collected local model and transmitting the packaged local model to an anonymous communication system, so that anonymization of the model is realized at a communication layer;
the terminal equipment of the Internet of things is used for collecting data in a local area network, then model training is carried out on a local data set to obtain local model update, anonymity of the model is achieved on a message layer by utilizing a ring signature technology, and finally the local model update is uploaded to a corresponding edge computing node.
2. A secure and fair federal learning privacy preserving aggregation method, comprising:
s1, initializing: terminal equipment C i Apply for Ring signing Key pair { sk i ,pk i Negotiating public key encryption algorithm g, hash function H and symmetric encryption algorithm E, edge node E j Apply for signing Key Pair { SK j ,PK j Establishing a P2P network among M edge nodes, and publishing a key pair { SK by a cloud server M ,PK M -and send the initialisation Model to all terminal devices;
s2, a local training stage: terminal equipment C i Receive the initialization Model andlocal training is completed to obtain local model update U i Updating the local model U i Encryption and attachment of a ring signature sigma i As model message M i Transmitting to an edge node;
s3, an anonymous communication system transmission stage: edge node E j Generating a message set G after receiving a certain number of model messages j ={M i |i∈S j S, where S j For E j The terminal equipment set under management forwards the message set to another edge node or cloud server according to the anonymous communication protocol;
s4, malicious model detection: model message M in received message set by cloud server i Performing ring signature verification, performing data dimension reduction on the model message, and performing malicious model detection through a homogeneous clustering technology;
s5, a global model updating stage: and (3) after all the malicious models detected in the step S4 are removed, the remaining benign models are aggregated, and global model updating is completed.
3. A secure and fair federal learning privacy preserving aggregation method according to claim 2, wherein S2 comprises:
s201, terminal equipment C i After the local model is trained, the model local update U is obtained i Simultaneously generating a random number N i According to public key PK of cloud server M Encryption to obtain ciphertext PK M (U i ||N i );
S202, terminal equipment C i Randomly picking a part of the participants' signature public key as a public key set R and calculating a ring signature
σ i =Sign ring (H(U i ||N i ) And send M i =(PK M (U i ||N i ),σ i ) As model messages to the corresponding edge computing node E j }。
4. A secure and fair federal learning privacy preserving aggregation method according to claim 2, wherein the anonymous communication protocol in S3 comprises:
s301, edge node E α Collecting a plurality of M i Combined into a message set G α ={M i |i∈S α Edge node E α Will G α With (G) α ,SK α (G α ||E β ) In the form of a) to another edge node E β
S302、E β Making a random selection, i.e. submitting to the cloud server with a probability of 1-p (G α ,SK γ (G α Server), and forwards the message with probability of p (G) α ,SK β (G α ||E γ ) For another edge node E) γ
S303、E γ A decision will be made with the same probability whether to submit the message (G α ,SK γ (G α Server) to the cloud Server, or continue forwarding messages;
s304, the cloud server receives the message (G) α ,SK j (G α Server), then for message set G α ={M i |i∈S α Each model message M in } i =(PK M (U i ||N i ),σ i ) Performing decryption calculation and verifying the ring signature sigma i To ensure the correctness of each model message M i Indeed by legal terminal equipment C i The generation; after the verification of the ring signature is passed, the random number N is checked i Ensure M i Not repeated messages.
5. A secure and fair federal learning privacy preserving aggregation method according to claim 3, wherein S4 comprises:
s401, the cloud server first receives all the received model messages M i =(PK M (U i ||N i ),σ i ) Performing ring signature verification to exclude all illegal model messages;
s402, constructing a legal model update set { U } i I is more than or equal to 1 and less than or equal to N, and is used for performing principal component analysis to realize data reductionDimension; and carrying out micro aggregation on the dimensionality reduced data to realize homogeneous clustering, and finally carrying out malicious model detection in each homogeneous class.
CN202311527204.6A 2023-11-16 2023-11-16 A secure and fair federated learning privacy-preserving aggregation system and method Pending CN117560159A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
CN202311527204.6A CN117560159A (en) 2023-11-16 2023-11-16 A secure and fair federated learning privacy-preserving aggregation system and method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
CN202311527204.6A CN117560159A (en) 2023-11-16 2023-11-16 A secure and fair federated learning privacy-preserving aggregation system and method

Publications (1)

Publication Number Publication Date
CN117560159A true CN117560159A (en) 2024-02-13

Family

ID=89810535

Family Applications (1)

Application Number Title Priority Date Filing Date
CN202311527204.6A Pending CN117560159A (en) 2023-11-16 2023-11-16 A secure and fair federated learning privacy-preserving aggregation system and method

Country Status (1)

Country Link
CN (1) CN117560159A (en)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119167432A (en) * 2024-11-21 2024-12-20 南昌大学 Cloud environment security monitoring method and device
CN119254400A (en) * 2024-09-12 2025-01-03 佳木斯大学 A hybrid privacy protection method for federated learning based on edge computing
CN119669888A (en) * 2024-10-22 2025-03-21 北京理工大学 A robust federated learning method and system based on secret sharing and privacy protection
CN120166125A (en) * 2025-03-05 2025-06-17 西安电子科技大学 A lightweight distributed model aggregation method for low-power chips in the Internet of Things

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119254400A (en) * 2024-09-12 2025-01-03 佳木斯大学 A hybrid privacy protection method for federated learning based on edge computing
CN119669888A (en) * 2024-10-22 2025-03-21 北京理工大学 A robust federated learning method and system based on secret sharing and privacy protection
CN119167432A (en) * 2024-11-21 2024-12-20 南昌大学 Cloud environment security monitoring method and device
CN120166125A (en) * 2025-03-05 2025-06-17 西安电子科技大学 A lightweight distributed model aggregation method for low-power chips in the Internet of Things

Similar Documents

Publication Publication Date Title
Kalapaaking et al. Blockchain-based federated learning with secure aggregation in trusted execution environment for internet-of-things
Hao et al. Efficient, private and robust federated learning
CN112861153B (en) Keyword searchable delayed encryption method and system
CN117560159A (en) A secure and fair federated learning privacy-preserving aggregation system and method
Xie et al. Verifiable federated learning with privacy-preserving data aggregation for consumer electronics
CN116049897A (en) Verifiable privacy protection federal learning method based on linear homomorphic hash and signcryption
Zhao et al. Efficient and privacy-preserving federated learning against poisoning adversaries
Eledlebi et al. Empirical studies of TESLA protocol: properties, implementations, and replacement of public cryptography using biometric authentication
Li et al. A Certificateless Pairing‐Free Authentication Scheme for Unmanned Aerial Vehicle Networks
Yang et al. Provably secure client‐server key management scheme in 5G networks
Aminanto et al. Multi-class intrusion detection using two-channel color mapping in IEEE 802.11 wireless network
Abidin On privacy-preserving biometric authentication
CN116957104A (en) Secure federal learning aggregation method, system, device and medium in wireless network
Mun et al. Emerging blockchain and reputation management in federated learning: Enhanced security and reliability for Internet of Vehicles (IoV)
Zhou et al. Group verifiable secure aggregate federated learning based on secret sharing
Kaushal et al. Securing the collective intelligence: a comprehensive review of federated learning security attacks and defensive strategies
Liu et al. Pivacy-preserving federated learning based on multi-key fully homomorphic encryption and trusted execution environment
Shayan Biscotti-a ledger for private and secure peer to peer machine learning
CN116707861B (en) A secure and robust feature combination method
Zhou et al. PPFLV: privacy-preserving federated learning with verifiability
Li et al. Efficient privacy aggregation method based on zero-knowledge proofs in federated learning
Wang et al. Dual-Server Privacy-Preserving Collaborative Deep Learning: A Round-Efficient, Dynamic and Lossless Approach
Wu et al. Vhfl: A cloud-edge model verification technique for hierarchical federated learning
Li et al. VMFL: A Verifiable Multi-Round Aggregation Scheme for Federated Learning in VANETs
Pichandi et al. Network security enhancement in data-driven intelligent architecture based on cloud IoT blockchain cryptanalysis

Legal Events

Date Code Title Description
PB01 Publication
PB01 Publication
SE01 Entry into force of request for substantive examination
SE01 Entry into force of request for substantive examination