TY - GEN
T1 - An Information Identification Method for Venture Firms Based on Frequent Itemset Discovery
AU - Cao, Ning
AU - Wang, Yansong
AU - Chen, Xiaoyu
AU - Zhou, Yulan
AU - Wu, Mingrui
AU - Li, Xiaofang
AU - Ding, Jianrui
AU - Zhu, Dongjie
N1 - Publisher Copyright:
© 2021, Springer Nature Switzerland AG.
PY - 2021
Y1 - 2021
N2 - In recent years, the emergence of a large number of venture firms has brought great profits to venture capital firms. However, it is not easy to identify venture firms with investment prospects. Therefore, based on frequent item sets, this paper mainly mines the enterprise text information to identify the venture enterprises with investment prospects. Firstly, we use TF-IDF algorithm to extract keywords from enterprise text; Secondly, the word2VEC model is used to vectorize the text keywords, and cosine similarity is calculated with the word vectors in the keyword database; Finally, we use the Apriori algorithm to find frequent item sets and generate association rules, complete vector weighting calculation of combination keywords, and finally retain the first three words or phrases with the highest weight as the identification keywords of the enterprise, thus determining whether the enterprise is a risk company with potential investment prospects. Experimental results show that the proposed method is effective.
AB - In recent years, the emergence of a large number of venture firms has brought great profits to venture capital firms. However, it is not easy to identify venture firms with investment prospects. Therefore, based on frequent item sets, this paper mainly mines the enterprise text information to identify the venture enterprises with investment prospects. Firstly, we use TF-IDF algorithm to extract keywords from enterprise text; Secondly, the word2VEC model is used to vectorize the text keywords, and cosine similarity is calculated with the word vectors in the keyword database; Finally, we use the Apriori algorithm to find frequent item sets and generate association rules, complete vector weighting calculation of combination keywords, and finally retain the first three words or phrases with the highest weight as the identification keywords of the enterprise, thus determining whether the enterprise is a risk company with potential investment prospects. Experimental results show that the proposed method is effective.
KW - Cosine similarity
KW - Discovery frequency term
KW - TF-IDF
KW - Venture firms identification
KW - Word2vec
UR - https://www.scopus.com/pages/publications/85111977868
U2 - 10.1007/978-3-030-78618-2_41
DO - 10.1007/978-3-030-78618-2_41
M3 - 会议稿件
AN - SCOPUS:85111977868
SN - 9783030786175
T3 - Communications in Computer and Information Science
SP - 496
EP - 509
BT - Advances in Artificial Intelligence and Security - 7th International Conference, ICAIS 2021, Proceedings
A2 - Sun, Xingming
A2 - Zhang, Xiaorui
A2 - Xia, Zhihua
A2 - Bertino, Elisa
PB - Springer Science and Business Media Deutschland GmbH
T2 - 7th International Conference on Artificial Intelligence and Security, ICAIS 2021
Y2 - 19 July 2021 through 23 July 2021
ER -