Skip to content
Features

Behind the benchmarks: SPEC, GFLOPS, MIPS et al

Ars discusses the terms and numbers behind benchmarking.

Jon Stokes – | 0
Story text

Behind the benchmarks: SPEC, GFLOPS, MIPS et al

What has the power to make or break a company or a career? What has the power to generate heated controversy, hard feelings, and bold accusations? A sex scandal? Litigation? Nope?try benchmarks. Ever since the days of multi-million dollar supercomputers and fat government contracts, benchmarks have played a critical role in the rise and fall of product lines and the people behind them. Though the Supercomputing Wars have just about gone the way of the Cold War of which they were so much a product, they have, in turn, given way to the modern-day Platform Wars. Advocates for platforms (hardware + OS) like Wintel, Mac, IRIX, and Alpha (Linux and NT, especially), fight tooth and nail in newsgroups and on web pages, wielding benchmarks with religious zeal in a pitched battle over whose platform is superior. As most of us know, things often get out of hand. Benchmarks are best used to help designers make informed design decisions and to help consumers make informed purchasing decisions. Using benchmarks to prove the God-given, moral superiority of your pet platform isn?t likely to win you many converts.

Platform zealots aren?t the only folks who use and abuse benchmarks; vendors have a vested interest in benchmarking, as well. In the hands of a vendor?s design team, benchmarks are essential tools for locating and eliminating performance bottlenecks and wasted work. In the hands of a vendor?s marketing department, however, benchmarks become tools of an entirely different sort. When it comes to evaluating most vendor-supplied benchmarks, the consumer would do well to adhere to maxim "caveat emptor." Companies have profits and reputations riding on these benchmarks. As a result, they have turned the practice of tweaking benchmarks into a high art.

Perhaps the most controversial, most often abused, and least-understood type of benchmarking is CPU benchmarking. CPU benchmarks come in a variety of shapes and sizes, but the SPEC CPU benchmarking suite is the most famous. People take SPEC benchmark results as holy writ. Other benchmark suites are popular, but impressive SPEC results are critical if a company is going to show how its CPU stacks up to the competition. But is SPEC, and other CPU benchmarks like it, really a good indicator of performance? What about GFLOPS and MIPS, the other two oft-quoted CPU ratings; how much do they tell us?

In this article, I?ll try to answer those questions by opening up the SPEC92 and SPEC95 CPU benchmark suites and seeing how they work. I?ll talk about their advantages and disadvantages, and the ways that companies cheat to jack up their numbers. But first, I?ll discuss MIPS, GFLOPS, and a few other commonly used performance metrics. In the end, I?ll give some conclusions and advice for the PC enthusiast community.

Note: When discussing CPU performance, a good starting place is Hennessy and Patterson?s seminal work on computer architecture, Computer Architecture: A Quantitative Approach. This book is the standard text in most computer architecture courses; when it comes to clarity of writing and depth of coverage, it is second to none. In the first chapter, Hennessy and Patterson lay out a set of principles and concepts for quantifying and reporting CPU performance. They also have information on the SPEC92 benchmark that I?ve not seen anywhere else. Thus, I?ll be basing the majority of this discussion on their work, while weaving in other sources as appropriate. Most of the other information used in this article is available from http://www.spec.org/.

MIPS and GFLOPS

Of all the misleading performance metrics out there, MIPS and GFLOPS ratings have got to be two of the most widespread. I myself have even been known to quote GFLOP numbers from vendors; when that?s all you have to go on, then you don?t have too much of a choice. Of the two, MIPS has to be the most worthless. (I?ve never been guilty of quoting MIPS.) For a given program, the MIPS (Millions of Instructions Per Second) rating is calculated by dividing the instruction count of the program by its execution time in milliseconds.

You don?t even have to think too hard to come up with a good criticism of MIPS. The instruction count for a particular program is completely dependent on the nature of the program and the instruction set architecture (ISA). For example, the "Hello World" program, when compiled for a CISC ISA, might use ten instructions; when compiled for a RISC ISA, it might use twenty. If the program takes the same amount of time to execute on both CPUs, the RISC chip will have MIPS rating twice as high as the CISC chip, even though both chips did the same amount of work in the same amount of time. In fact, not only does MIPS vary drastically between ISAs, it also varies between different programs on the same computer.

As bad as MIPS is for rating performance, GFLOPS isn?t much better. To measure GFLOPS (billion Floating Point Operations Per Second), you divide the number of floating point operations in a program by the execution time in milliseconds. One problem with this approach is the fact that the number of floating point ops varies from program to program. Two programs, one that?s 80% floating-point ops and one that?s 20% floating-point ops, both of which take the same amount of time to execute, will have different GFLOPS ratings.

Another, even bigger problem with GFLOPS is that the not all the same floating point instructions are implemented on all machines. One machine may use two floating-point ops to perform a particular task, while another machine may use only one. If the task is completed in the same amount of time on both machines, the one that used two ops to do it will have a higher GFLOPS rating.

In short, neither GFLOPS nor MIPS provides a reliable metric for gauging performance. The next time you see a MIPS or GFLOPS rating, notice the source?I?ll 99% guarantee you it?s a vendor. The reason for this is twofold. First, a vendor is the only one who?s really going to put in the time and effort that it takes to count up the instruction mix for a program and do all the other stuff you have to do to assign a MIPS or FLOPS rating. Second, vendors are the only people who benefit from such a rating. Most consumers don?t know enough about a vendor?s architecture to be able to determine which floating-point ops are available on it versus which FP ops are available on competing architectures. Even if a consumer did have this info, the vendor never divulges what program was used for the rating or what the instruction mix for that program was, so it wouldn?t be of any use.

Fortunately, there are other options available.

Using real-world apps

When measuring computer performance, all that ultimately matters is that the programs you use every day run as fast as possible. Thus, one of the most popular, and best, ways to benchmark a system is by running a real program on it and see how long it takes the program to complete a task. While this approach is excellent for benchmarking an entire system (RAM, disk subsystem, motherboard, etc.), it doesn?t provide a reliable measure of CPU performance. Whenever you fire up Quake2 and run a timedemo, the final FPS rating depends on memory latency, hard disk performance, video card performance, and a host of other factors. You also have to take into account the fact that a multitasking OS will be running a number of other processes besides your program; not only do those other processes take up time, but so does the context switch required to move from one program to the next. (You have to save the register contents, the stack pointer, etc.)

The amount of CPU time a program actually gets vs. the amount of time it spends waiting for I/O or in various system calls varies from program to program. The CPU time actually spent in the program is called user time. The CPU time spent in the operating system or in another program is called system time. If most of a program?s total execution time is system time, then you have to ask yourself what you?re benchmarking: the device drivers and I/O subsystem, or the CPU. Ideally, you want to run a benchmark that spends as much time as possible in user time.

The upshot of all this is that, while running a suite of CPU-intensive programs and putting a stopwatch on them may be a good source of info about the performance of a complete platform, it doesn?t tell you too much about CPU performance in specific. Application performance becomes an even less realistic CPU benchmark if you?re comparing two different platforms. The application may have been extensively tuned for maximum performance on one platform, and merely ported to run on another. Thus, running Photoshop on a Mac and then running it on a PC doesn?t tell you too much about how the G3 stacks up against the PII. All it tells you is how well Photoshop (and only Photoshop) runs on a Mac as compared to a PC. Photoshop is a excellent example of this point. Caesar is currently running benchmarks on the G3 and comparing them to various Intel processors. Within the first few tests, he hit snags related to memory, OS architecture, and more.

In light of the above discussion, it doesn?t sound like real applications are a very good way to measure CPU performance. If all you?re interested in is CPU performance, real application benchmarks are not the best route. But if you?re looking to benchmark a complete system (which is the only sane thing to try to benchmark, anyway), then you can?t do better than real-world apps.

Using benchmarking programs

Since real-world applications don?t allow you to zone in on a particular component of a platform, an alternative is to use a program designed specifically for benchmarking CPU performance. There are a number of such programs out there, and they fall into different categories. Hennessy and Patterson identify three different categories of benchmark programs (apart from real-world applications).

The first type of program is called a kernel. A kernel is just a small, CPU-intensive portion of a real program that?s intended to be run on its own or as part of a benchmarking suite. The idea behind kernels is that if you take out all the I/O requests and other system calls then you can maximize the amount of user time the program gets. In the design phase, kernels can be great tools for profiling the performance of specific parts of a machine and chasing down inefficiencies in a design. The problem is when companies employ them for marketing purposes.

When kernels are used as part of a standard benchmarking suite?and the SPEC suite contains a number of kernels?they become targets for compiler optimization. Compiler writers can aggressively optimize for a few specific kernels and jack up the overall benchmark scores significantly. Hennessy and Patterson report that on some systems, compiler optimizations can increase the performance of SPEC?s nasa7 kernel by 2.1 times. When this sort of thing happens, you?re not benchmarking the hardware, but the compiler. Many compilers include benchmark-specific optimizations that never get used in a real-world applications; their only purpose is to increase performance on one specific benchmark.

Actually, when you consider how dependent modern CPUs are on compiler technology, it doesn?t sound so bad to say you?re "benchmarking the compiler." In upcoming next-generation architectures like Intel?s IA-64, the line between the compiler and the hardware is intentionally blurred. So compiler optimization isn?t really "cheating", unless the optimization doesn?t increase performance in any real-world applications. In fact, SPEC95 claims to benchmark the CPU, memory architecture, and the compiler. These three components are so interdependent that it really is impossible to benchmark one of them individually.

The second type of benchmarking program Hennessy and Patterson identify is what they call a toy benchmark. Toy benchmarks are small programs like Quicksort or the Sieve of Eratosthenes that people cook up to compile and run on any computer. Such benchmarks don?t really measure much of anything, and in the opinion of the authors, "the best use of such programs is beginning programming assignments."

The third type of benchmark is the synthetic benchmark. A synthetic benchmark, like Dhrystone or Whetstone, tries to determine the average instruction mix of a typical program so that it can replicate that mix and run those types of instructions on the CPU. Synthetic benchmarks, by definition, aren?t real programs, and their code doesn?t compute anything anyone would necessarily ever want to compute. As such, the results they yield may or may not be useful in predicting the performance of real programs. There are a whole host of problems with synthetic benchmarks; the topic is worthy of an article in and of itself. Hennessy and Patterson consider these benchmarks the least useful of all the types discussed.

Using real-world apps cont.

Actually, I will cover one criticism of synthetic benchmarks, because it can apply to all of the benchmarks I’ve just mentioned in this section. That criticism is this: a benchmark application?s sole purpose in life is to produce benchmark results, results which are, by definition, not real-world results. What do I mean? Take the popular synthetic benchmark suite, WinBench. WinBench et al are fine if you want a quick-and-dirty comparison that can give you a rough estimate of performance, but you need to ask yourself if Winbench’s CPUMark 99 scores are anything you really care about. Face it, all Winbench’s CPUMark 99 scores tell you is how well your machine runs CPUMark 99?something which may or may not be indicative of how your machine will run the apps you actually use. Do you do your spreadsheets in CPUMark 99? How about your word-processing? Are you a member of a Winbench clan? If the answer to the above questions is "no," then maybe Winbench’s CPUMark 99 scores aren?t as important as Ziff-Davis would have you think.

SPEC

Quite often, you can get around the problems inherent in just using one benchmarking program by using a collection of them to exercise different parts of a system. This is the idea behind the benchmarking suite. The most popular benchmarking suite for gauging the performance of CPUs is the SPEC suite. From the SPEC website:

The basic SPEC methodology is to provide the benchmarker with a standardized suite of source code based upon existing applications that has already been ported to a wide variety of platforms by its membership. The benchmarker then takes this source code, compiles it for the system in question and then can tune the system for the best results. The use of already accepted and ported source code greatly reduces the problem of making apples-to-oranges comparisons.

The SPEC suites consist of a set of floating-point-intensive and integer-intensive kernels and programs. Vendors get the source for the programs and compile it themselves before running the benchmarks. The vendor is required to submit two sets of scores: the baseline performance measurement and the optimized performance measurement. The baseline performance measurement restricts the vendor to a single set of flags and a single compiler for all the benchmarks in the same language. The SPEC92 suite let vendors use a home-grown compilers with a long and impenetrable list of arcane options to generate the baseline score; such compiler gymnastics can have a considerable effect on the baseline score. The SPEC95 suite specifies a set of tools with which to build and run the benchmark programs, and they limit the number of compiler flags to four, not counting portability flags.

In SPEC92, the optimized performance measurement allowed vendors to go wild, using benchmark-specific compiler flags that, if used on a normal program, would generate either illegal code or code that?s considerably slower. In some cases, the difference between the optimized and unoptimized scores could be substantial. With the release of SPEC95, some efforts have been made to reign vendors in. SPEC has added some rules aimed at eliminating the use of flags that would normally produce unsafe code, or that a vendor wouldn?t endorse for use with a regular application. SPEC now imposes the following conventions on the use of compiler optimizations:

  • Hardware and software used to run the CINT95/CFP95 benchmarks must provide a suitable environment for running typical C or FORTRAN programs.
  • Optimizations must generate correct code for a class of programs, where the class of programs must be larger than a single SPEC benchmark or SPEC benchmark suite. This also applies to "assertion flags" that may be used for non-base/"aggressive compilation" measurements (see section 2.2.4 p2).
  • Optimizations must improve performance for a class of programs where the class of programs must be larger than a single SPEC benchmark or SPEC benchmark suite.
  • The vendor encourages the implementation for general use.
  • The implementation is generally available, documented and supported by the providing vendor.

(WARNING: this following paragraph may make your eyes glaze over. It makes mine glaze over and I?m the one who wrote it. If you don?t follow this, don?t sweat it.)

To summarize performance, a SPEC suite uses the geometric mean of the normalized execution times of its constituent programs. What this means is that first, the execution time for each program is normalized to the base execution time for that same program on a reference machine. For SPEC92, the reference machine is a VAX-11/780; for SPEC95, the reference machine is a SPARCstation 10/40 (40MHz SuperSPARC with no L2 cache). To illustrate the use of normalized ratio, consider this example: if the gcc test took 10 seconds on the reference machine and 1 second on the benchmarked machine, then gcc has a normalized execution time of 0.1 seconds on the benchmarked machine. The n normalized execution times are then multiplied together, and the nth root of the result is taken; this gives the normalized geometric mean of the execution times.

What does this mean in English? It means that, in the end, what matters is not how many seconds it took for a particular program to run, but how many times faster that program ran on the benchmarked machine than it ran on the reference machine. In our gcc example, the benchmarked machine ran the gcc program ten times faster than the reference machine. To calculate the final score, you figure up how many times faster each program ran on the test machine than it did on the reference machine, and you take the geometric mean of those ratios. I?m not going to try and explain the geometric mean here; just take it on faith that it works better than taking the arithmetic mean (or average) of the ratios.

There are a number of pros and cons to using the geometric mean as a summary of performance. One of the pros of using the geometric mean is that it’s independent of the actual running times of the individual programs?remember, all execution times are normalized. Also, it doesn?t matter which machine is used as a reference, since the scores are all relative. One of the drawbacks to using the normalized geometric mean is the fact that it encourages the vendors to focus on improving the scores of the programs that are the easiest to crack. If a program runs in two seconds and, after optimizing some aspect of the design (the compiler or the hardware), the vendor gets it to run in one second, they get the same increase in the overall score as they would have if they had gotten a program that runs in 1000 seconds to run in 500 seconds. So substantial increases in performance are rewarded the same as insubstantial ones if the ratio of increase is the same.

Numbers

An actual SPEC95 report of the results of a benchmarking run consists of a whole slew of numbers and graphs. Here are the eight most important numbers, ripped right out of the SPEC95 Q&A:

CINT95:

  • SPECint95: The geometric mean of eight normalized ratios (one for each integer benchmark) when compiled with aggressive optimization for each benchmark.
  • SPECint_base95: The geometric mean of eight normalized ratios (one for each integer benchmark) when compiled with conservative optimization for each benchmark.
  • SPECint_rate95: The geometric mean of eight normalized throughput ratios (one for each integer benchmark) when compiled with aggressive optimization for each benchmark.
  • SPECint_rate_base95: The geometric mean of eight normalized throughput ratios (one for each integer benchmark) when compiled with conservative optimization for each benchmark.

CFP95:

  • SPECfp95: The geometric mean of 10 normalized ratios (one for each floating point benchmark) when compiled with aggressive optimization for each benchmark.
  • SPECfp_base95: The geometric mean of 10 normalized ratios (one for each floating point benchmark) when compiled with conservative optimization for each benchmark.
  • SPECfp_rate95: The geometric mean of 10 normalized throughput ratios (one for each floating point benchmark) when compiled with aggressive optimization for each benchmark.
  • SPECfp_rate_base95: The geometric mean of 10 normalized throughput ratios (one for each floating point benchmark) when compiled with conservative optimization for each benchmark.

The ratios for each of the benchmarks are calculated based on a SPEC-determined reference time and the run time of the benchmark.

SPECint95, SPECint_base95, SPECfp95, and SPECfp_base95 measure speed, or how long the computer takes to complete a single task. The other "rate" numbers measure throughput, or the ability of the computer to handle multiple tasks.

Now let?s take a look at what?s actually in the SPEC92 and SPEC95 benchmarks.

Table 1: SPECint92

Benchmark Source Lines of code Description
espresso C 13,500 Minimizes Boolean functions.
li C 7,4131 A lisp interpreter written in C that solves the 9-queens problem.
eqntott C 3,376 Translates a Boolean equation into a truth table.
compress C 1,503 Performs data compression on a 1-MB file using Lempel-Ziv coding.
sc C 8,116 Performs computations within a UNIX-spreadsheet.
gcc C 83,589 Consists of the GNU C compiler converting preprocessed files into optimized Sun-3 machine code.

 

from Computer Architecture: a Quantitative Approach, by John L. Hennessy and David A. Patterson, p. 22

 

Table 2: SPECfp92

Benchmark Source Lines of code Description
spice2g6 FORTRAN 18,476 Circuit simulation package that simulates a small circuit.
doduc FORTRAN 5,334 A Monte Carlo simulation of a nuclear reactor component.
mdljdp2 FORTRAN 4,458 A chemical application that solves equations of motion for a model of 500 atoms. This is similar to modeling the structure of liquid argon.
wave5 FORTRAN 7,628 A two-dimensional electromagnetic particle-in-cell simulation used to study various plasma phenomena. Solves equations of motion on a mesh involving 500,000 particles on 50,000 grid points for 5 time steps.
tomcatv FORTRAN 195 A mesh generation program, which is highly vectorizable.
ora FORTRAN 535 Traces rays through optical systems of spherical and plane surfaces.
mdljsp2 FORTRAN 3,885 Same as mdljdp2, but single precision.
alvinn C 272 Simulates training of a neural network. Uses single precision.
ear C 4,483 An inner ear model that filters and detects various sounds and generates speech signals. Uses single precision.
swm256 FORTRAN 487 A shallow water model that solves shallow water equations using finite difference equations with a 256 x 256 grid. Uses single precision.
su2cor FORTRAN 487 Computes masses of elementary particles from Quark-Gluon theory.
hydro2d FORTRAN 4,461 An astrophysics application program that solves hydrodynamical Navier Stokes equations to compute galactical jets.
nasa7 FORTRAN 1,204 Seven kernels d matrix manipulation, FFTs, Gaussian elimination, vortices creation.
fpppp FORTRAN 2,718 A quantum chemistry application program used to calculate two electron integral derivatives.

 

from Computer Architecture: a Quantitative Approach, by John L. Hennessy and David A. Patterson, p. 22

 

Table 3: CINT95

Benchmark Description
099.go An internationally ranked go-playing program.
124.m88ksim A chip simulator for the Motorola 88100 microprocessor..
126.gcc Based on the GNU C compiler version 2.5.3.
129.compress A in-memory version of the common UNIX utility.
130.li Xlisp interpreter.
132.ijpeg Image compression/decompression on in-memory images.
134.perl An interpreter for the Perl language.
147.vortex An object-oriented database.

 

from http://www.spec.org/osg/cpu95/CINT95/

 

Table 4: CFP95

Benchmark Description
101.tomcatv Vectorized mesh generation.
102.swim Shallow water equations.
103.su2cor Monte-Carlo method.
104.hydro2d Navier Stokes equations.
107.mgrid 3D potential field.
110.applu Image compression/decompression on in-memory images.
125.turb3d Turbulence modeling.
141.apsi Weather prediction.
145.fpppp From Gaussian series of quantum chemistry benchmarks.
146.wave5 Maxwell’s equations.

 

from http://www.spec.org/osg/cpu95/CFP95/

Looking at the above tables, you may have noticed that some of the programs and kernels perform functions that you use on a regular basis, like image compression, and functions you might never use, like weather prediction. The SPEC suite is definitely biased towards hardcore technical computing, and with good reason. The kinds of people who do weather prediction and SPICE simulations are the kinds of people who are trying to decide which machine to spend that $60,000 in grant money on. They?re not the people who are choosing between a K6-2 and a PII. This fact also explains why the SPEC results are normalized to a base machine that most of us won?t ever own: the SPARCstation 10/40.

The above comments lead me to a criticism that applies to a number of benchmarking suites. Remember how I mentioned that compiler optimizations can make significant differences in the performance of an application or benchmark? Well, clean, optimized, thoughtfully written code can also really set an application apart from its peers. (Actually, a marketing juggernaut and a virtual monopoly can also set an application apart from its peers.) So if you?re going to use an application for benchmarking, it?s best to use one that you actually use every day. Similarly, if you?re going to use an application suite for benchmarking, it too needs to be made up of apps that you use every day. Who cares how fast a certain 3D renderer runs if it isn?t the one you?re using? Different renderers scale differently across different CPUs, so you need to benchmark with the particular renderer you use. The same applies to all other software. How it was coded, and how that code was optimized and compiled, makes worlds of difference as to how it runs and how scalable it is. The bottom line is that the performance of programs that you don?t ever run (except as benchmarks) may or may not be indicative of the performance of programs that you actually do run, and the performance of programs you actually run is the only performance you should care about.

Note: I acknowledge that many of the kernels in the SPEC benchmarking suite are pieces of code that show up regularly in common applications (Gaussian elimination, FFT, etc.). This fact still doesn?t significantly detract from my overall point. Said code is always part of a larger application that?s subject to all of the criticisms I just outlined.

In light of the above observations, you can see that while the SPEC benchmarks are thorough and informative, aside from the occasional usenet platform flame-fest, they?re of limited use to the PC hardware enthusiast community. Besides, we hardcore PC types like to do our own benchmarking, but the SPEC benchmarks are too expensive. A SPEC95 CD-ROM will set you back about $600 if you?re a new customer, so you won?t be seeing any SPEC results coming out of the Damage Labs anytime soon.

Conclusions

Conclusions

No matter what benchmark you ultimately decide on, there?s no getting around the fact that real-world applications are the best tools for measuring overall performance. The main question the consumer should take into consideration is: how fast will this system run my most demanding applications? Benchmarks can be rigged. but applications can?t… usually. (I do know of one instance where a game was rigged to give better framerates when using Intel?s new SSE. It’s obvious from these screenshots that the SSE-enabled version has a lower poly count.) Certain benchmarks may claim to measure only one or two components in a system, but the reality is that when you benchmark, you benchmark a whole machine/hardware, OS, and all.

Another thing that it always helps to remember when factoring benchmark scores into an upgrade decision is that what matters is how fast the apps you use now run on the machine, not how fast applications expected to ship with the next version of the OS are projected to run once the developers add support for the latest whiz-bang tech (SSE, MMX, etc.). If your favorite app doesn’t yet support Intel’s new UltraSHAFT Retro-Fro Hair Extensions, then said extensions aren’t doing you much good, are they? If you have to wait six months for support to come out so you can see that 10% performance gain on odd days of the week when the moon is full, then you have to ask yourself if you shouldn’t have just waited and bought a chip with a higher MHz rating.

Product cycles fly by too fast to hold off on a sorely needed upgrade until corporate promises materialize. If you’ve got to upgrade (like I will when Q3A comes out), you just have to hold your nose and do it at some point, and you have to do it based on what’s out there on the day you’re going to upgrade. Sure, you can wait about a month if you know there’s a major price cut coming, or if you know an actual release date for a new product that will drive prices down, but when you start waiting two or three months, then why not wait two or three more? There’s always something better coming down the pike. This is where good benchmarks can help: they tell you what product to buy today, when "today" is the day you’re ordering the upgrade.

So in the end, there’s only one way to test performance. First, make sure all your drivers are the latest and greatest that are currently available. Then, fire up your favorite app and put the stopwatch on it to see how long it takes it to do its thing. I know that people will continue using synthetic benchmarks, kernels, and suites as a matter of practicality; most real-world apps don?t generate charts, graphs, and numerical scores to tell you how we?ll they?re performing on your system. In fact, we at Ars Technica have been known to use synthetic benchmarks ourselves, and we?ll probably do so again (unless companies like Adobe start hooking us up with free software). But the reality is that you can?t find a better benchmark than a real app. If you’re banking on any other type of benchmark to tell you how to spend your $$$$, then caveat emptor.

0 Comments

Comments are closed.