Disclosure of Invention
The present application has been made in view of the above problems, and has as its object to provide a picture recognition method, a computer device, and a computer-readable storage medium, which overcome or at least partially solve the above problems.
According to an aspect of the present application, there is provided a picture recognition method including:
Acquiring content information, wherein the content information comprises a plurality of information elements, and the plurality of information elements comprise at least one picture information element;
extracting characteristic data of information elements in the content information;
Determining difference data between the picture information element and other information elements according to the characteristic data;
and determining the characteristic type of the content information according to the characteristic data and the difference data.
Optionally, the determining the feature type of the content information according to the feature data and the difference data includes:
taking the characteristic data of the information element as input, and determining the characteristic value of the information element according to a characteristic value identification model;
and determining the characteristic type of the content information according to the characteristic value and the difference data.
Optionally, before the determining the feature value of the information element according to the feature value recognition model with the feature data of the information element as input, the method further comprises:
And training the characteristic value recognition model by adopting the information element sample and the characteristic type of the corresponding mark.
Optionally, the content information includes a text information element, the feature data includes description information, and determining difference data between the picture information element and other information elements according to the feature data includes:
and determining difference data between the picture information element and the text information element according to the description information of the picture information element and the description information of the text information element.
Optionally, determining difference data between the picture information element and other information elements according to the feature data includes:
and comparing the characteristic data among the picture information elements to obtain difference data among the picture information elements.
Optionally, determining difference data between the picture information element and other information elements according to the feature data includes:
clustering a plurality of picture information elements according to the characteristic data;
and calculating difference data among the picture clusters as the difference data among the picture information elements.
Optionally, before the extracting the characteristic data of the information element in the content information, the method further includes:
searching for associated information of the content information, wherein the associated information comprises at least one of the following: picture information element, text information element, video information element;
The association information is added to the content information.
Optionally, the association information includes a picture information element, and before the searching for the association information of the content information, the method further includes:
And determining that the number of picture information elements in the content information does not meet preset requirements.
Optionally, the content information includes at least one of comment content information and video content information.
Optionally, when the content information is video content information, before the extracting the feature data of the information element in the content information, the method further includes:
and extracting a plurality of video frames in the video content information as the picture information element.
According to another aspect of the present application, there is provided a picture recognition apparatus including:
an information acquisition module for acquiring content information, the content information including a plurality of information elements including at least one picture information element;
the data extraction module is used for extracting characteristic data of information elements in the content information;
The difference determining module is used for determining difference data between the picture information element and other information elements according to the characteristic data;
and the type determining module is used for determining the characteristic type of the content information according to the characteristic data and the difference data.
Optionally, the type determining module includes:
the characteristic value determining submodule is used for taking characteristic data of the information element as input and determining the characteristic value of the information element according to a characteristic value identification model;
and the type determining sub-module is used for determining the characteristic type of the content information according to the characteristic value and the difference data.
Optionally, the apparatus further comprises:
And the training module is used for training the characteristic value recognition model by adopting the information element sample and the characteristic type of the corresponding mark before the characteristic value of the information element is determined according to the characteristic value recognition model by taking the characteristic data of the information element as input.
Optionally, the content information includes a text information element, the feature data includes description information, and the difference determining module includes:
And the difference determining sub-module is used for determining difference data between the picture information element and the text information element according to the description information of the picture information element and the description information of the text information element.
Optionally, the difference determining module includes:
and the comparison sub-module is used for comparing the characteristic data among the picture information elements to obtain the difference data among the picture information elements.
Optionally, the difference determining module includes:
the clustering sub-module is used for clustering the plurality of picture information elements according to the characteristic data;
and the calculating sub-module is used for calculating difference data among the picture clusters and taking the difference data as the difference data among the picture information elements.
Optionally, the apparatus further comprises:
The searching module is used for searching the associated information of the content information before the characteristic data of the information element in the content information is extracted, and the associated information comprises at least one of the following: picture information element, text information element, video information element;
And the adding module is used for adding the association information into the content information.
Optionally, the association information includes a picture information element, and the apparatus further includes:
The determining module is used for determining that the number of the picture information elements in the content information does not meet the preset requirement before the associated information of the content information is searched.
Optionally, the content information includes at least one of comment content information and video content information.
Optionally, when the content information is video content information, the apparatus further includes:
And the video frame extraction module is used for extracting a plurality of video frames in the video content information as the picture information elements before the characteristic data of the information elements in the content information are extracted.
According to another aspect of the present application there is provided a computer device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, the processor implementing a method as one or more of the above when executing the computer program.
According to another aspect of the present application, there is provided a computer readable storage medium having stored thereon a computer program which when executed by a processor performs a method as one or more of the above.
According to the embodiment of the application, the content information comprises a plurality of information elements, the plurality of information elements comprise at least one picture information element, the characteristic data of the information elements in the content information are extracted, the difference data between the picture information elements and other information elements are determined according to the characteristic data, the characteristic type of the content information is determined according to the characteristic data and the difference data, the analysis of the dimension of the difference between pictures is introduced, the characteristic type can be analyzed from the whole of the content information, the problem of lower accuracy when the single picture is relied on to analyze the characteristics is avoided, and the accuracy of determining the characteristic type of the content information is improved.
Further, the related information is added to the content information by searching the related information of the content information, so that the obtained content information is supplemented, the problem that information elements in the content information are insufficient is solved, the problem that the mode cannot be implemented due to the insufficient information elements is solved, or the dimension of performing feature analysis on the content information can be increased by increasing the number of the information elements, and the accuracy of the feature analysis is further provided.
The foregoing description is only an overview of the present application, and is intended to be implemented in accordance with the teachings of the present application in order that the same may be more clearly understood and to make the same and other objects, features and advantages of the present application more readily apparent.
Detailed Description
Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be embodied in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
For a better understanding of the present application, the following description is given to illustrate the concepts related to the present application to those skilled in the art:
The content information includes information in the form of pictures, text, video, audio, etc., and in the present application, the content information is composed of a plurality of information elements and at least one picture information element is included in the plurality of information elements. The information elements may include information elements in the form of pictures, text, video, audio, etc., or any other suitable form, as embodiments of the application are not limited in this regard. The content information may have a single form of information element or may have multiple forms of information element. One or more of the various information elements may be present.
For example, in an e-commerce transaction, a commodity evaluation submitted for a transaction commodity belongs to a piece of content information, which may include a picture information element, a text information element, etc., or the content of a question-answer area in a commodity detail page also belongs to a piece of content information. Or in the video network platform, videos such as film and television drama belong to content information, the videos consist of video frame sequences, video frames of the videos can be used as picture information elements, and audio of the videos can be used as audio information elements.
The feature data includes various data used to characterize the information elements, such as, for example, for picture information elements, using a spatial vector model to represent each picture information element, vectorizing each picture information element, or any other suitable feature data, as embodiments of the present application are not limited in this respect. The description information of the information element may also be regarded as a kind of characteristic data, for example, the subject word of the text information element is a kind of description information, the subject of the text may be obtained using TextCNN (text convolutional neural network) model, the subject word may describe a certain characteristic of the text, and thus the subject word may be regarded as a kind of characteristic data of the text information element. Any suitable descriptive information may be included in particular, and embodiments of the application are not limited in this regard.
The difference data between the information elements is used to characterize the magnitude of the difference between the information elements, and may be obtained by comparing the feature data of the information elements, or any other suitable manner, which is not limited by the embodiments of the present application. The difference data may be difference data between two information elements or difference data between a plurality of information elements, where the difference data between a plurality of information elements may include a sum value or an average value of the difference data between every two information elements, which is not limited in the embodiment of the present application.
In order to overcome the problem that malicious pictures are difficult to identify depending on a single picture information element, the difference data to be determined is between the picture information element and other picture information elements or information elements in non-picture form. The greater the difference between a picture information element and other information elements, the greater the likelihood that the picture information element is a malicious picture and, in turn, the greater the likelihood that the content information is malicious information. Conversely, the smaller the difference between a picture information element and other information elements, the less likely the picture information element is a malicious picture, and consequently the less likely the content information is malicious information.
For example, for the difference data between the picture information elements, the picture information elements may be represented by vectors, the difference data may be obtained by cosine similarity calculation in one way, and the similarity of the two picture information elements may be evaluated by calculating the cosine values of the vectors of the two picture information elements, that is, the difference data of the two picture information elements may be obtained; in another mode, pearson correlation coefficient (Pearson correlation coefficient) can be adopted to calculate to obtain difference data, and the correlation between the vector of two picture information elements and the quotient of the standard deviation can be evaluated by calculating the covariance between the vector of the two picture information elements, so that the difference data of the two picture information elements are obtained, and in particular, the difference data can be obtained in any applicable mode, and the embodiment of the application is not limited to the above.
The content information may be classified into a feature type according to the feature data and the difference data of the information element, the feature type is used for characterizing a feature of a certain dimension of the content information, for example, for commodity evaluation, the feature type may be classified into advertisement evaluation and normal evaluation, or may be classified into malicious evaluation and non-malicious evaluation, or may be classified into various kinds of serious violations, slight violations, suspected violations, non-violations, and the like; for the video, the feature types may be classified into a replaced advertisement and an un-replaced advertisement, or may be classified into a face change or an un-face change, or may be classified into a picture change or an un-picture change, or may be classified into a plurality of types such as a video change or an un-video change, and specifically may include any suitable feature type, which is not limited in this embodiment of the present application.
In an alternative embodiment of the application, in order to determine the feature type of the content information, the degree of the information element in the content information in such feature dimension may be determined first and recorded as a feature value. The characteristic value of the information element is used to characterize the value of the degree of the information element on the corresponding characteristic. Taking the characteristic of malicious degree as an example, in a numerical value interval of 0-1, the larger the numerical value is, the more serious the malicious degree is, and the malicious degree of a certain information element is 0.6, which is more serious. It will be appreciated that the characteristic values may also be represented by text or symbols, for example, the level of maliciousness of the information element is ten (e.g., level 1-10 corresponds to the preceding value interval of 0-1).
For this purpose, a feature value recognition model is used to determine the feature data of the information element, and the feature value recognition model is input with the feature data of the information element, and the model can output the feature value of the information element through calculation. The feature value of the information element can be output because the feature value identification model adopts a supervised learning mode, and the model is trained according to a large number of information element samples of the marked feature type, so that the feature value identification model can evaluate the feature value of the information element.
For example, a large number of picture information elements of advertisement evaluation and normal evaluation are collected, the picture information elements are marked by adopting two characteristic types of the advertisement evaluation and the normal evaluation, a two-class model is obtained by training in a supervised learning mode, the two-class model can attribute the picture information elements to the two characteristic types of the advertisement evaluation or the normal evaluation, but the picture information elements output by the two-class model are not required to be classified to be of a certain characteristic type, and parameters which can represent the degree of the picture information elements on the characteristic type in the two-class model are output as the characteristic values of the picture information elements.
In an alternative embodiment of the application, when the content information comprises a text information element, in order to determine the difference data between the picture information element and the text information element, it is necessary to obtain first the description information of the information element, which also belongs to a kind of characteristic data.
For example, the description information of the text information element may be obtained as the description information by using TextCNN model or TF-IDF (term frequency-inverse frequency) algorithm, etc. The description information of the picture information element can be obtained by adopting a Mask R-CNN (Mask Region-based Convolutional Neural Network, mask-based convolutional neural network) model or a picture scene recognition model and other modes, and the main target or scene included in the picture information element can be used as the description information. The description information of the information element may be obtained in any suitable manner, which is not limited in this embodiment of the present application.
In an alternative embodiment of the application, the content information has information associated with it in an application system or data system, denoted as association information. The association information includes at least one of: picture information elements, text information elements, video information elements, or any other suitable form, to which embodiments of the application are not limited.
For example, for content information such as commodity comments, the commodity for which the commodity comments are made will have detailed information of the commodity in the electronic commerce platform, and the detailed information is related information of the commodity comments. Including text describing the commodity, a display diagram of the commodity, and even a demonstration video of the commodity use process. And storing the commodity ID (identification) correspondingly when the commodity comments are stored, inquiring according to the commodity ID, and searching the detailed information of the commodity and extracting the required information elements from the detailed information. Any suitable searching method may be specifically adopted, and the embodiment of the present application is not limited thereto.
In an alternative embodiment of the application, when the content information is video content information, i.e. the acquired content information is data in a video format, the video is composed of a sequence of video frames, which may be regarded as picture information elements.
For example, the scheme of the application can be used for processing video by an infringer when an advertiser implants an advertised commodity in a certain scene in the video in a tamper-detection application scene, and tampering the advertised commodity into other commodities in the video at partial time. Or in the application scene of video advertisement detection, namely detecting whether the video frame embedded with the advertisement exists in the video. Or in an application scenario for personal privacy detection, i.e. detecting whether there are video frames of content related to personal privacy in the video.
In one implementation manner for the application scenario, the following description is given: the method comprises the steps of obtaining video frames at different times in a scene in a video, taking the obtained video frames as picture information elements, extracting characteristic data of the picture information elements, determining difference data between the picture information elements, and then determining whether the video is a tampered video or a video implanted with advertisements or a video related to personal privacy according to the characteristic data and the difference data.
In another implementation manner for the application scenario described above: the content information comprises a video to be detected and a sample video frame (such as a tampered video frame, or a video frame implanted with an advertisement, or a video frame related to personal privacy, etc.), the video frame sequence to be detected is input, tampering or advertisement detection is carried out, characteristic data of the video frame to be detected and the sample video frame are extracted, difference data between the video frame to be detected and the sample video frame are determined, and then whether the video is the tampered video, the video implanted with the advertisement, or the video related to personal privacy is determined according to the characteristic data and the difference data. The sample video frames may have positive samples or negative samples. For example, a positive sample is an untampered video frame, then a negative sample is a tampered video frame; positive samples are video frames with an advertisement embedded, then negative samples are video frames without an advertisement embedded.
According to an embodiment of the present application, when the evaluation content includes malicious pictures, for example, the malicious pictures include various hidden advertisement information, and the manner of identifying malicious based on a single picture has a problem of low accuracy. As shown in a schematic diagram of a picture recognition process in fig. 1, the present application provides a picture recognition mechanism, by acquiring content information, where the content information includes a plurality of information elements, where the plurality of information elements includes at least one picture information element, extracting feature data of the information elements in the content information, determining difference data between the picture information elements and other information elements according to the feature data, determining a feature type of the content information according to the feature data and the difference data, and introducing analysis of a dimension of difference between pictures, so that the feature type can be analyzed from the whole of the content information, thereby avoiding a problem of lower accuracy when a single picture is relied on to analyze features, and improving accuracy of determining the feature type of the content information. The application is applicable but not limited to the above application scenario.
Referring to fig. 2, a flowchart of an embodiment of a picture recognition method according to a first embodiment of the present application is shown, and the method may specifically include the following steps:
Step 101, obtaining content information, wherein the content information comprises a plurality of information elements, and the plurality of information elements comprise at least one picture information element.
In order to solve the problem that malicious pictures are difficult to identify, the embodiment of the application firstly obtains the content information of the feature type to be determined, wherein the content information consists of a plurality of information elements and comprises at least one picture information element.
For example, in an electronic commerce transaction, when a user submits a commodity evaluation to a transaction commodity, the commodity evaluation is acquired from a data system of the electronic commerce and analyzed before the commodity evaluation is displayed on a comment page, so that the commodity evaluation with problems is prevented from being published on the comment page. Or in the video network platform, after the user uploads the self-made video, before the video is provided for other users to watch, the video is acquired from the data system of the video network platform for analysis, and of course, for the video content information, the video frames need to be extracted from the video as picture information elements.
And 102, extracting characteristic data of information elements in the content information.
In the embodiment of the application, the implementation modes for extracting the characteristic data of the information element in the content information comprise various modes, for the picture information element, the picture information element can be vectorized, the obtained vector is used as the characteristic data of the picture information element, and if the difference data between the picture information element and the text information element is required to be determined later, the picture information element can be subjected to target recognition or scene recognition to obtain the description information such as the main target or the scene where the picture is located and the like, and the description information is used as the characteristic data. For the text information element, description information such as a subject word of a text can be identified as feature data. Any suitable extraction method may be specifically adopted, and the embodiment of the present application is not limited thereto.
For example, as shown in fig. 3, a schematic diagram of a risk identification process of a commodity comment includes 1 text information element and 5 picture information elements. And carrying out vectorization representation on each picture information element to obtain a vector corresponding to each picture information element, wherein the vector is used as characteristic data of the picture information element. The subject word of the text information element is identified as characteristic data of the text information element.
And step 103, determining difference data between the picture information element and other information elements according to the characteristic data.
In the embodiment of the application, the difference data between the picture information element and other information elements can be determined according to the characteristic data of the information element. When the content information includes a plurality of picture information elements, the picture information elements and other information elements include picture information elements and other picture information elements, and may also include picture information elements and other information elements in a non-picture form. When the content information includes one picture information element, the picture information element and other information elements include picture information elements and other information elements in non-picture form.
In the embodiment of the present application, the implementation manner of determining the difference data between the picture information elements and other information elements may include various ways, for example, comparing feature data between the picture information elements to obtain the difference data between the picture information elements; or clustering a plurality of picture information elements according to the characteristic data, and calculating difference data among the picture clusters to serve as the difference data among the picture information elements; or determining difference data between the picture information element and the text information element according to the description information of the picture information element and the description information of the text information element, or any other suitable implementation manner, which is not limited by the embodiment of the present application.
For example, as shown in fig. 3, the comparison is performed between every two of the 5 picture information elements, and the difference degree value, that is, the difference data, between every two of the 5 picture information elements is obtained by comparing the feature data corresponding to the picture information elements, and the obtained 5 difference data may be directly used, or 1 or more difference data with the largest difference in the 5 difference data may be selected for use, or a mean value of the 5 difference data may be calculated to obtain a total difference data for use, or any other applicable manner may be used.
And 104, determining the characteristic type of the content information according to the characteristic data and the difference data.
In the embodiment of the application, the content information can be assigned to a certain feature type according to the feature data of the information element and the difference data obtained in the last step. Specifically, the determination may be performed according to the feature data of all the information elements in the content information and the difference data obtained in the previous step, or may be determined according to the feature data of part of the information elements in the content information and the difference data obtained in the previous step, which is not limited in the embodiment of the present application.
In the embodiment of the present application, the implementation manner of determining the feature type of the content information may include various implementations, for example, taking the feature data of the information element as input, determining the feature value of the information element according to the feature value recognition model, determining the feature type of the content information according to the feature value and the difference data, or any other applicable implementation manner, which is not limited in this embodiment of the present application.
For example, as shown in fig. 3, feature data of 1 text information element in a commodity comment, namely, a subject word, is input into a risk value recognition model for a text to obtain a risk value of the text information element, weighted average calculation is performed according to the risk value and a difference degree value obtained in the last step to obtain a comprehensive risk value, and if the comprehensive risk value exceeds a preset threshold, it is determined that the commodity comment belongs to a feature type with risk.
According to the embodiment of the application, the content information comprises a plurality of information elements, the plurality of information elements comprise at least one picture information element, the characteristic data of the information elements in the content information are extracted, the difference data between the picture information elements and other information elements are determined according to the characteristic data, the characteristic type of the content information is determined according to the characteristic data and the difference data, the analysis of the dimension of the difference between pictures is introduced, the characteristic type can be analyzed from the whole of the content information, the problem of lower accuracy when the single picture is relied on to analyze the characteristics is avoided, and the accuracy of determining the characteristic type of the content information is improved.
Referring to fig. 4, a flowchart of an embodiment of a picture recognition method according to a second embodiment of the present application is shown, and the method specifically may include the following steps:
step 201, obtaining content information, the content information comprising a plurality of information elements, the plurality of information elements comprising at least one picture information element.
In the embodiments of the present application, the specific implementation manner of this step may be referred to the description in the foregoing embodiments, which is not repeated herein.
In the embodiment of the present application, optionally, the content information includes at least one of comment content information and video content information. For example, in an e-commerce transaction, a commodity evaluation submitted by a user for a commodity belongs to one type of comment content information. In the video network platform, self-made video uploaded by a user belongs to video content information. The comment content information or the video content information can be arbitrarily applied, and the embodiment of the application is not limited to the comment content information or the video content information.
Step 202, searching association information of the content information, wherein the association information comprises at least one of the following: picture information element, text information element, video information element.
In the embodiment of the application, the associated information of the content information can be searched before the characteristic data of the information element in the content information is extracted. For example, for the commodity comment, the detailed information of the commodity aimed at by the commodity comment is searched, the picture in the detailed information is determined as the associated information, and specifically, the searching can be performed according to the actual requirement, which is not limited by the embodiment of the present application.
In the embodiment of the present application, optionally, in some implementation scenarios, difference data between the picture information elements is required, and the association information to be searched includes the picture information elements. Before searching the associated information of the content information, the method may further include: and determining that the number of picture information elements in the content information does not meet the preset requirement. Specifically, the number of preset requirements can be set according to actual needs, which is not limited in the embodiment of the present application.
For example, the commodity comment includes only 1 picture information element, and the number of the preset requirements is 2, and when the number of the picture information elements in the commodity comment is determined to be not in accordance with the preset requirements, the picture information elements in the content information are indicated to be insufficient. And searching pictures in the detailed information of the commodity according to the commodity comment to be used as associated information and supplementing the associated information into the content information.
And step 203, adding the association information to the content information.
In the embodiment of the application, the searched associated information is added into the content information to realize the supplement of the acquired content information, solve the problem of insufficient information elements in the content information, and overcome the problem that the mode cannot be implemented due to insufficient information elements, or the dimension of the feature analysis of the content information can be increased by increasing the number of the information elements, so that the accuracy of the feature analysis is further provided.
And 204, extracting and extracting characteristic data of information elements in the content information.
In the embodiments of the present application, the specific implementation manner of this step may be referred to the description in the foregoing embodiments, which is not repeated herein.
Step 205, clustering a plurality of picture information elements according to the feature data.
In the embodiment of the application, when calculating the difference data among the plurality of picture information elements, a clustering mode can be adopted for the plurality of picture information elements, and the process of dividing the plurality of picture information elements into a plurality of classes consisting of similar picture information elements is called clustering.
In the embodiment of the application, the characteristic data of the plurality of picture information elements are clustered, namely the plurality of picture information elements are clustered. Clusters generated by clustering are a set of picture information elements that are similar to picture information elements in the same cluster, and picture information elements of one cluster are noted as a picture cluster, unlike picture information elements in other clusters. For example, 5 picture information elements in a commodity comment are clustered, wherein 3 picture information elements are assigned to one picture cluster and 2 picture information elements are assigned to another picture cluster.
Step 206, calculating difference data between the picture clusters as the difference data between the picture information elements.
In the embodiment of the application, after clustering, difference data among the picture clusters can be calculated and used as the difference data among the picture information elements. The implementation manner of calculating the difference data between the picture clusters may include various ways, for example, calculating the difference data between each picture information element in one picture cluster and each picture information element in another picture cluster to obtain a plurality of difference data, and then calculating an average value of the plurality of difference data to obtain the difference data between the two picture clusters. Any suitable computing means may be included, and embodiments of the present application are not limited in this regard.
In an embodiment of the present application, optionally, the content information includes a text information element, the feature data includes description information, and an implementation manner of determining difference data between the picture information element and other information elements according to the feature data may include: and determining difference data between the picture information element and the text information element according to the description information of the picture information element and the description information of the text information element.
And comparing the description information of the picture information element with the description information of the text information element to obtain difference data between the picture information element and the text information element. The differences between the descriptive information may characterize differences between the picture information element and the text information element. For example, text vectorization is performed on the two pieces of description information, the distance between vectors corresponding to the two pieces of description information is calculated to obtain difference data, the description information of the picture information element is "sock", and when the description information of the text information element is "cotton cloth", the obtained difference data is smaller than that when the description information of the text information element is "plastic".
Step 207, taking the characteristic data of the information element as input, and determining the characteristic value of the information element according to a characteristic value recognition model.
In the embodiment of the application, when the characteristic value of the information element is determined by adopting the characteristic value identification model, the input data is the characteristic data of the information element, and the characteristic value identification model can output the characteristic value of the information element after calculation.
In an embodiment of the present application, optionally, before the feature data of the information element is taken as an input, and the feature value of the information element is determined according to a feature value recognition model, the method may further include: and training the characteristic value recognition model by adopting the information element sample and the characteristic type of the corresponding mark.
The eigenvalue recognition model needs to be trained to accurately recognize the eigenvalue of the information element. And collecting a large number of information element samples, marking the feature types corresponding to the information elements, inputting the information element samples into a feature value identification model, and continuously learning and updating parameters in the feature value identification model until the required performance is achieved.
And step 208, determining the characteristic type of the content information according to the characteristic value and the difference data.
In the embodiment of the application, after the characteristic value is obtained by adopting the characteristic value identification model, the characteristic value and the difference data are integrated to determine the characteristic type of the content information. For example, after the weighted average calculation is performed on the feature value and the difference data to obtain a comprehensive value, if the comprehensive value exceeds a preset threshold, the content information is determined to be of one feature type, and if the comprehensive value does not exceed the preset threshold, the content information is determined to be of another feature type, and any suitable implementation manner may be specifically adopted, which is not limited in the embodiment of the present application.
According to the embodiment of the application, the content information comprises a plurality of information elements, the plurality of information elements comprise at least one picture information element, the characteristic data of the information elements in the content information are extracted, the plurality of picture information elements are clustered according to the characteristic data, the difference data among the picture clusters is calculated and is used as the difference data among the picture information elements, the characteristic data of the information elements is used as input, the characteristic value of the information elements is determined according to a characteristic value identification model, the characteristic type of the content information is determined according to the characteristic value and the difference data, the characteristic type is analyzed from the whole content information, the problem of lower accuracy when the characteristic is analyzed according to a single picture is avoided, and the accuracy of determining the characteristic type of the content information is improved.
Further, the related information is added to the content information by searching the related information of the content information, so that the obtained content information is supplemented, the problem that information elements in the content information are insufficient is solved, the problem that the mode cannot be implemented due to the insufficient information elements is solved, or the dimension of performing feature analysis on the content information can be increased by increasing the number of the information elements, and the accuracy of the feature analysis is further provided.
Referring to fig. 5, a block diagram illustrating an embodiment of a picture recognition apparatus according to a third embodiment of the present application may specifically include:
an information acquisition module 301, configured to acquire content information, where the content information includes a plurality of information elements, and the plurality of information elements includes at least one picture information element;
A data extraction module 302, configured to extract feature data of information elements in the content information;
A difference determining module 303, configured to determine difference data between the picture information element and other information elements according to the feature data;
And a type determining module 304, configured to determine a feature type of the content information according to the feature data and the difference data.
In an embodiment of the present application, optionally, the type determining module includes:
the characteristic value determining submodule is used for taking characteristic data of the information element as input and determining the characteristic value of the information element according to a characteristic value identification model;
and the type determining sub-module is used for determining the characteristic type of the content information according to the characteristic value and the difference data.
In an embodiment of the present application, optionally, the apparatus further includes:
And the training module is used for training the characteristic value recognition model by adopting the information element sample and the characteristic type of the corresponding mark before the characteristic value of the information element is determined according to the characteristic value recognition model by taking the characteristic data of the information element as input.
In an embodiment of the present application, optionally, the content information includes a text information element, the feature data includes description information, and the variance determining module includes:
And the difference determining sub-module is used for determining difference data between the picture information element and the text information element according to the description information of the picture information element and the description information of the text information element.
In an embodiment of the present application, optionally, the difference determining module includes:
and the comparison sub-module is used for comparing the characteristic data among the picture information elements to obtain the difference data among the picture information elements.
In an embodiment of the present application, optionally, the difference determining module includes:
the clustering sub-module is used for clustering the plurality of picture information elements according to the characteristic data;
and the calculating sub-module is used for calculating difference data among the picture clusters and taking the difference data as the difference data among the picture information elements.
In an embodiment of the present application, optionally, the apparatus further includes:
The searching module is used for searching the associated information of the content information before the characteristic data of the information element in the content information is extracted, and the associated information comprises at least one of the following: picture information element, text information element, video information element;
And the adding module is used for adding the association information into the content information.
In an embodiment of the present application, optionally, the association information includes a picture information element, and the apparatus further includes:
The determining module is used for determining that the number of the picture information elements in the content information does not meet the preset requirement before the associated information of the content information is searched.
In the embodiment of the present application, optionally, the content information includes at least one of comment content information and video content information.
In an embodiment of the present application, optionally, when the content information is video content information, the apparatus further includes:
And the video frame extraction module is used for extracting a plurality of video frames in the video content information as the picture information elements before the characteristic data of the information elements in the content information are extracted.
According to the embodiment of the application, the content information comprises a plurality of information elements, the plurality of information elements comprise at least one picture information element, the characteristic data of the information elements in the content information are extracted, the difference data between the picture information elements and other information elements are determined according to the characteristic data, the characteristic type of the content information is determined according to the characteristic data and the difference data, the analysis of the dimension of the difference between pictures is introduced, the characteristic type can be analyzed from the whole of the content information, the problem of lower accuracy when the single picture is relied on to analyze the characteristics is avoided, and the accuracy of determining the characteristic type of the content information is improved.
For the device embodiments, since they are substantially similar to the method embodiments, the description is relatively simple, and reference is made to the description of the method embodiments for relevant points.
Embodiments of the present disclosure may be implemented as a system configured as desired using any suitable hardware, firmware, software, or any combination thereof. Fig. 6 schematically illustrates an example system (or apparatus) 700 that may be used to implement various embodiments described in this disclosure.
For one embodiment, FIG. 6 illustrates an exemplary system 700 having one or more processors 702, a system control module (chipset) 704 coupled to at least one of the processor(s) 702, a system memory 706 coupled to the system control module 704, a non-volatile memory (NVM)/storage 708 coupled to the system control module 704, one or more input/output devices 710 coupled to the system control module 704, and a network interface 712 coupled to the system control module 706.
The processor 702 may include one or more single-core or multi-core processors, and the processor 702 may include any combination of general-purpose or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, system 700 can function as a browser as described in embodiments of the present application.
In some embodiments, the system 700 can include one or more computer-readable media (e.g., system memory 706 or NVM/storage 708) having instructions and one or more processors 702 coupled with the one or more computer-readable media and configured to execute the instructions to implement the modules to perform the actions described in this disclosure.
For one embodiment, the system control module 704 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 702 and/or any suitable device or component in communication with the system control module 704.
The system control module 704 may include a memory controller module to provide an interface to the system memory 706. The memory controller modules may be hardware modules, software modules, and/or firmware modules.
The system memory 706 may be used to load and store data and/or instructions for the system 700, for example. For one embodiment, system memory 706 may comprise any suitable volatile memory, such as, for example, a suitable DRAM. In some embodiments, the system memory 706 may comprise a double data rate type four synchronous dynamic random access memory (DDR 4 SDRAM).
For one embodiment, system control module 704 may include one or more input/output controllers to provide interfaces to NVM/storage 708 and input/output device(s) 710.
For example, NVM/storage 708 may be used to store data and/or instructions. NVM/storage 708 may include any suitable nonvolatile memory (e.g., flash memory) and/or may include any suitable nonvolatile storage device(s) (e.g., one or more Hard Disk Drives (HDDs), one or more Compact Disc (CD) drives, and/or one or more Digital Versatile Disc (DVD) drives).
NVM/storage 708 may include a storage resource that is physically part of the device on which system 700 is installed, or it may be accessed by the device without being part of the device. For example, NVM/storage 708 may be accessed over a network via input/output device(s) 710.
Input/output device(s) 710 may provide an interface for system 700 to communicate with any other suitable device, input/output device 710 may include communication components, audio components, sensor components, and the like. Network interface 712 may provide an interface for system 700 to communicate over one or more networks, and system 700 may communicate wirelessly with one or more components of a wireless network according to any of one or more wireless network standards and/or protocols, such as accessing a wireless network based on a communication standard, such as WiFi,2G,3G,4G, or 5G, or a combination thereof.
For one embodiment, at least one of the processor(s) 702 may be packaged together with logic of one or more controllers (e.g., memory controller modules) of the system control module 704. For one embodiment, at least one of the processor(s) 702 may be packaged together with logic of one or more controllers of the system control module 704 to form a System In Package (SiP). For one embodiment, at least one of the processor(s) 702 may be integrated on the same die with logic of one or more controllers of the system control module 704. For one embodiment, at least one of the processor(s) 702 may be integrated on the same die with logic of one or more controllers of the system control module 704 to form a system on chip (SoC).
In various embodiments, system 700 may be, but is not limited to being: a browser, workstation, desktop computing device, or mobile computing device (e.g., a laptop computing device, handheld computing device, tablet, netbook, etc.). In various embodiments, system 700 may have more or fewer components and/or different architectures. For example, in some embodiments, system 700 includes one or more cameras, keyboards, liquid Crystal Display (LCD) screens (including touch screen displays), non-volatile memory ports, multiple antennas, graphics chips, application Specific Integrated Circuits (ASICs), and speakers.
Wherein if the display comprises a touch panel, the display screen may be implemented as a touch screen display to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensor may sense not only the boundary of a touch or slide action, but also the duration and pressure associated with the touch or slide operation.
The embodiment of the application also provides a non-volatile readable storage medium, wherein one or more modules (programs) are stored in the storage medium, and when the one or more modules are applied to a terminal device, the terminal device can be caused to execute instructions (instructions) of each method step in the embodiment of the application.
In one example, a computer device is provided comprising a memory, a processor and a computer program stored on the memory and executable on the processor, characterized in that the processor implements a method according to an embodiment of the application when executing the computer program.
There is also provided in one example a computer readable storage medium having stored thereon a computer program, characterized in that the program when executed by a processor implements a method as in one or more of the embodiments of the application.
The embodiment of the application discloses a picture identification method and a picture identification device, and example 1 comprises a picture identification method, comprising the following steps:
Acquiring content information, wherein the content information comprises a plurality of information elements, and the plurality of information elements comprise at least one picture information element;
extracting characteristic data of information elements in the content information;
Determining difference data between the picture information element and other information elements according to the characteristic data;
and determining the characteristic type of the content information according to the characteristic data and the difference data.
Example 2 may include the method of example 1, wherein the determining the feature type of the content information from the feature data and the difference data comprises:
taking the characteristic data of the information element as input, and determining the characteristic value of the information element according to a characteristic value identification model;
and determining the characteristic type of the content information according to the characteristic value and the difference data.
Example 3 may include the method of example 1 and/or example 2, wherein, prior to the determining the characteristic value of the information element from the characteristic value recognition model with the characteristic data of the information element as input, the method further comprises:
And training the characteristic value recognition model by adopting the information element sample and the characteristic type of the corresponding mark.
Example 4 may include the method of one or more of examples 1-3, wherein the content information includes a text information element, the feature data includes descriptive information, and determining difference data between the picture information element and other information elements based on the feature data includes:
and determining difference data between the picture information element and the text information element according to the description information of the picture information element and the description information of the text information element.
Example 5 may include the method of one or more of examples 1-4, wherein the determining difference data between the picture information element and other information elements from the feature data comprises:
and comparing the characteristic data among the picture information elements to obtain difference data among the picture information elements.
Example 6 may include the method of one or more of examples 1-5, wherein the determining difference data between the picture information element and other information elements from the feature data comprises:
clustering a plurality of picture information elements according to the characteristic data;
and calculating difference data among the picture clusters as the difference data among the picture information elements.
Example 7 may include the method of one or more of examples 1-6, wherein prior to the extracting the feature data of the information element in the content information, the method further comprises:
searching for associated information of the content information, wherein the associated information comprises at least one of the following: picture information element, text information element, video information element;
The association information is added to the content information.
Example 8 may include the method of one or more of examples 1-7, wherein the association information includes a picture information element, the method further comprising, prior to the locating the association information of the content information:
And determining that the number of picture information elements in the content information does not meet preset requirements.
Example 9 may include the method of one or more of examples 1-8, wherein the content information includes at least one of comment content information, video content information.
Example 10 may include the method of one or more of examples 1-6, wherein, when the content information is video content information, prior to the extracting the feature data of the information element in the content information, the method further comprises:
and extracting a plurality of video frames in the video content information as the picture information element.
Example 11 includes a picture recognition apparatus, comprising:
an information acquisition module for acquiring content information, the content information including a plurality of information elements including at least one picture information element;
the data extraction module is used for extracting characteristic data of information elements in the content information;
The difference determining module is used for determining difference data between the picture information element and other information elements according to the characteristic data;
and the type determining module is used for determining the characteristic type of the content information according to the characteristic data and the difference data.
Example 12 may include the apparatus of example 11, wherein the type determination module comprises:
the characteristic value determining submodule is used for taking characteristic data of the information element as input and determining the characteristic value of the information element according to a characteristic value identification model;
and the type determining sub-module is used for determining the characteristic type of the content information according to the characteristic value and the difference data.
Example 13 may include the apparatus of example 11 and/or example 12, wherein the apparatus further comprises:
And the training module is used for training the characteristic value recognition model by adopting the information element sample and the characteristic type of the corresponding mark before the characteristic value of the information element is determined according to the characteristic value recognition model by taking the characteristic data of the information element as input.
Example 14 may include the apparatus of one or more of examples 11-13, wherein the content information includes a text information element, the feature data includes descriptive information, and the variance determining module includes:
And the difference determining sub-module is used for determining difference data between the picture information element and the text information element according to the description information of the picture information element and the description information of the text information element.
Example 15 may include the apparatus of one or more of examples 11-14, wherein the variance determination module comprises:
and the comparison sub-module is used for comparing the characteristic data among the picture information elements to obtain the difference data among the picture information elements.
Example 16 may include the apparatus of one or more of examples 11-15, wherein the variance determination module comprises:
the clustering sub-module is used for clustering the plurality of picture information elements according to the characteristic data;
and the calculating sub-module is used for calculating difference data among the picture clusters and taking the difference data as the difference data among the picture information elements.
Example 17 may include the apparatus of one or more of examples 11-16, wherein the apparatus further comprises:
The searching module is used for searching the associated information of the content information before the characteristic data of the information element in the content information is extracted, and the associated information comprises at least one of the following: picture information element, text information element, video information element;
And the adding module is used for adding the association information into the content information.
Example 18 may include the apparatus of one or more of examples 11-17, wherein the association information includes a picture information element, the apparatus further comprising:
The determining module is used for determining that the number of the picture information elements in the content information does not meet the preset requirement before the associated information of the content information is searched.
Example 19 may include the apparatus of one or more of examples 11-18, wherein the content information includes at least one of comment content information, video content information.
Example 20 may include the apparatus of one or more of examples 11-19, wherein when the content information is video content information, the apparatus further comprises:
And the video frame extraction module is used for extracting a plurality of video frames in the video content information as the picture information elements before the characteristic data of the information elements in the content information are extracted.
Example 21 includes a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implementing the method as in one or more of examples 1-10 when the computer program is executed.
Example 22 includes a computer-readable storage medium having stored thereon a computer program that, when executed by a processor, performs a method as in one or more of examples 1-10.
While certain embodiments have been illustrated and described for purposes of description, various alternative, and/or equivalent embodiments, or implementations calculated to achieve the same purposes are shown and described without departing from the scope of the embodiments of the present application. This disclosure is intended to cover any adaptations or variations of the embodiments discussed herein. It is manifestly, therefore, that the embodiments described herein are limited only by the claims and the equivalents thereof.