码迷,mamicode.com
首页 > 其他好文 > 详细

pyspark GBT

时间:2018-04-17 19:53:15      阅读:265      评论:0      收藏:0      [点我收藏+]

标签:basic   tap   col   ons   ssi   tmp   port   nbsp   collect   

from pyspark import SparkContext

from pyspark import SparkConf

from pyspark.mllib.regression import LabeledPoint

from pyspark.mllib.tree import GradientBoostedTrees

 

string_test = ‘pyspark_test‘

conf = SparkConf().setAppName(string_test).setMaster(‘yarn‘)

sc = SparkContext(conf=conf)

hdfs_data = sc.textFile("hdfs://dap/basicdata/dianzhang/phoenix/tmp_dianzhang_train_type/ds=2018-04-15/type=R1_C1/000000_0")

a = hdfs_data.map(lambda x:x.split(‘\t‘))

b = a.map(lambda x:LabeledPoint(x.pop(-1), x[3:]))

‘‘‘

弯路

c = b.collect()

model = GradientBoostedTrees.trainClassifier(sc.parallelize(c), {}, numIterations=10)

‘‘‘

model = GradientBoostedTrees.trainClassifier(b, {}, numIterations=10)

 

pyspark GBT

标签:basic   tap   col   ons   ssi   tmp   port   nbsp   collect   

原文地址:https://www.cnblogs.com/kayy/p/8868500.html

(0)
(0)
   
举报
评论 一句话评论(0
登录后才能评论!
© 2014 mamicode.com 版权所有  联系我们:gaon5@hotmail.com
迷上了代码!