我的hadoop程序如下所示。我放了一些相关的代码片段。我通过的论点,读大数据在主要是真的。总的来说,“大数据工作是印刷出来的”。但当涉及rowpremap类中的map方法时,大数据的值是它的初始化值false。不知道为什么会这样。我错过什么了吗?当我在一台独立的机器上运行这个程序时,它就起作用了,但当我在hadoop集群上运行这个程序时,它就不起作用了。作业由jobcontrol处理。是有线的东西吗?
公共类uvdriver扩展配置的实现工具{
public static class RowMPreMap extends MapReduceBase implements
Mapper<LongWritable, Text, Text, Text> {
private Text keyText = new Text();
private Text valText = new Text();
public void map(LongWritable key, Text value,
OutputCollector<Text, Text> output, Reporter reporter)
throws IOException {
// Input: (lineNo, lineContent)
// Split each line using seperator based on the dataset.
String line[] = null;
if (Settings.BIG_DATA)
line = value.toString().split("::");
else
line = value.toString().split("\\s");
keyText.set(line[0]);
valText.set(line[1] + "," + line[2]);
// Output: (userid, "movieid,rating")
output.collect(keyText, valText);
}
}
public static class Settings {
public static boolean BIG_DATA = false;
public static int noOfUsers = 0;
public static int noOfMovies = 0;
public static final int noOfCommonFeatures = 10;
public static final int noOfIterationsRequired = 3;
public static final float INITIAL_VALUE = 0.1f;
public static final String NORMALIZE_DATA_PATH_TEMP = "normalize_temp";
public static final String NORMALIZE_DATA_PATH = "normalize";
public static String INPUT_PATH = "input";
public static String OUTPUT_PATH = "output";
public static String TEMP_PATH = "temp";
}
public static class Constants {
public static final int BIG_DATA_USERS = 71567;
public static final int BIG_DATA_MOVIES = 10681;
public static final int SMALL_DATA_USERS = 943;
public static final int SMALL_DATA_MOVIES = 1682;
public static final int M_Matrix = 1;
public static final int U_Matrix = 2;
public static final int V_Matrix = 3;
}
public int run(String[] args) throws Exception {
// 1. Pre-process the data.
// a) Normalize
// 2. Initialize the U, V Matrices
// a) Initialize U Matrix
// b) Initialize V Matrix
// 3. Iterate to update U and V.
// Write Job details for each of the above steps.
Settings.INPUT_PATH = args[0];
Settings.OUTPUT_PATH = args[1];
Settings.TEMP_PATH = args[2];
Settings.BIG_DATA = Boolean.parseBoolean(args[3]);
if (Settings.BIG_DATA) {
System.out.println("Working on BIG DATA.");
Settings.noOfUsers = Constants.BIG_DATA_USERS;
Settings.noOfMovies = Constants.BIG_DATA_MOVIES;
} else {
System.out.println("Working on Small DATA.");
Settings.noOfUsers = Constants.SMALL_DATA_USERS;
Settings.noOfMovies = Constants.SMALL_DATA_MOVIES;
}
// some code here
handleRun(control);
return 0;
}
public static void main(String args[]) throws Exception {
System.out.println("Program started");
if (args.length != 4) {
System.err
.println("Usage: UVDriver <input path> <output path> <fs path>");
System.exit(-1);
}
Configuration configuration = new Configuration();
String[] otherArgs = new GenericOptionsParser(configuration, args)
.getRemainingArgs();
ToolRunner.run(new UVDriver(), otherArgs);
System.out.println("Program complete.");
System.exit(0);
}
}
作业控制。
public static class JobRunner implements Runnable {
private JobControl control;
public JobRunner(JobControl _control) {
this.control = _control;
}
public void run() {
this.control.run();
}
}
public static void handleRun(JobControl control)
throws InterruptedException {
JobRunner runner = new JobRunner(control);
Thread t = new Thread(runner);
t.start();
int i = 0;
while (!control.allFinished()) {
if (i % 20 == 0) {
System.out
.println(new Date().toString() + ": Still running...");
System.out.println("Running jobs: "
+ control.getRunningJobs().toString());
System.out.println("Waiting jobs: "
+ control.getWaitingJobs().toString());
System.out.println("Successful jobs: "
+ control.getSuccessfulJobs().toString());
}
Thread.sleep(1000);
i++;
}
if (control.getFailedJobs() != null) {
System.out.println("Failed jobs: "
+ control.getFailedJobs().toString());
}
}
1条答案
按热度按时间ekqde3dh1#
这是行不通的,因为静态修饰符的作用域不跨越jvm的多个示例(更不用说网络)
Map任务总是在单独的jvm中运行,即使它是在工具运行程序的本地运行的。Map器类仅使用类名示例化,无法访问您在工具运行器中设置的信息。
这就是配置框架存在的原因之一。