You Can get it GuideLine
You Can get it Here Click 下载 to Download it
Train Data Format
| type | Text |
|---|---|
| entertainment | 台媒预测周冬雨金马奖封后,大气的倪妮却佳作难出 |
| food | 农村就是好,能吃到纯天然无添加的野生蜂蜜,营养又健康 |
| fashion | 14款知性美装,时尚惊艳搁浅的阳光轻熟的优雅 |
| etc. | etc. |
- 文本Text去除停用词 cut Text and drop stop_words
- 使用Tf-idf计算词向量值,建立词向量表示矩阵 construct Word Vector Matrix by calculating word vector number using Tf-idf
It is worth mentioning that: if there are too many features, sparse matrix can help you save a lot of memory - 建立模型 creat Model
I tried many Models(svm、cnn、logisticRegression,etc.) on this task butFINALLYI chosed MultinomialNB just because it cost least time and got high precision recall and F1-score - 预测结果 predict
| Test/Train % | avg-precision | avg-recall | avg-F1-score | time-cost(s) |
|---|---|---|---|---|
| 100% | 93% | 93% | 93% | 8s |
| 75% | 93% | 93% | 93% | 7s |
| 50% | 93% | 93% | 93% | 7s |
| 25% | 93% | 93% | 93% | 7s |
| 0% | 77% | 77% | 77% | 7s |
This means that training characteristics and model stability are high.This model has a good generalization,but in terms of accuracy, I think there is still room for improvement by improving the features
更多的结果表示方法我已经放在Text_Classifier用于文本预测分类中,可以直接pip安装使用。
You can get other represent of prediction at Text_Classifier

