Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
igodard
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
igodard
10y ago
ConAsm (model-dependent assembler) for "Silver" model: F("fact") %0; sub(b0 %0, 1) %1; gtrsb(b0 %1, 1) %2, retnfl(b2 %0), inner("fact$1_1", b1 %1, b2 %0); L("fact$1
32.
▲
by
igodard
10y ago
Perhaps you are thinking that the FPGA is a product? It's not; it's an RTL validator. Moving chip RTL from FPGA to product silicon is a well understood step that is almost routine in the industry. Time was you would do initial RTL
33.
▲
by
igodard
10y ago
Evolutionary development you can start at once, because you are building on what you had before; think an x86 generation. Evolution works if you already dominate a market and only need to run a little faster than your competitor. Evolution
34.
▲
by
igodard
10y ago
Exactly. As a new entrant, it is impossible for us to immediately enter the mass markets dominated by the majors. Consequently we adopted a strategy of targeting an increasing set of niche markets that have been poorly server by the majors
35.
▲
by
igodard
11y ago
(Mill team) The streams are not streams of instruction; they are streams of half-instructions. The Mill is a wide-issue machine, like a VLIW or EPIC; each instruction can contain many operations, which all issue together. Each instruction i
36.
▲
by
igodard
12y ago
Some responses to comments here: We do not have an FPGA implementation, although we are working on it. The reason getting to product is slow is that we are by choice a bootstrap startup, with three full-timers and a dozen or so part-timers.
37.
▲
by
igodard
12y ago
Thank you.
38.
▲
by
igodard
12y ago
The original of this discussion was a blog post on kevmod.com. I posted the following comment on that blog, repeated here verbatim as possibly of general interest: ++++++++++++++++++++++++++++++++++++++++++++++++++++ Your skepticism is comp
39.
▲
by
igodard
12y ago
Mill uses optimistic concurrency, similar to the IBM and Intel versions. From that you can build mutexes if you are willing to put up with the drawbacks of locking.
40.
▲
by
igodard
12y ago
Mill multicore has fully sequentially consistent cache coherency; there are no barrier operations. Sorry, how it's done is still NYF (Not Yet Filed). We expect a talk on the subject this fall.
41.
▲
by
igodard
12y ago
It's actually rather easy to allocate belt positions, because the belt is a FIFO and allocates them itself :-) However, the scheduler must track lifetimes and make sure that nothing still live falls off the end. This is also (not quite
42.
▲
by
igodard
12y ago
If you watch the Belt talk on our site, and know how a modern OOO machine works on the inside, then you will recognize that the Belt is a forwarding network, sometimes also called a bypass. There is no RAM, "S" or otherwise, no ge
43.
▲
by
igodard
12y ago
No funding by giants; not a public company. The SEC rules prevent us from talking further (we're too busy to go to jail for breaking the securities regs), but if you are interested in the business side of the company then you can sign
44.
▲
by
igodard
13y ago
My first compiler (still in use) was for the Burroughs B6500 mainframe in 1970. During my brief and inglorious college career I did not take a CS class. In fact, there were no CS classes. The college didn't even own a computer. Yes, th
45.
▲
by
igodard
13y ago
Not quite: only scheduling and binary creation is done at install time. Instruction selection and optimization is done during compilation in the usual way.
46.
▲
by
igodard
13y ago
You then assume your own conclusion when you ignore pipelining. If instructions are executed in sequence indian-file then necessarily none will be faster than any other. The traditional rule-of-thumb is that programs have an ILP of two. The
47.
▲
by
igodard
13y ago
It would be nice if there were a single number that could be justified by measurement, but there's no hardware yet to measure and there would not be a single number even if the hardware existed. That's because there's not jus
48.
▲
by
igodard
13y ago
Contribute it.
49.
▲
by
igodard
13y ago
Exactly. Most will specialize at install-time. Code without an install step will specialize at first execution; the IR is designed for specialize speed. The specializer can also be run free-standing as well, to create ROMS and where install
50.
▲
by
igodard
13y ago
#1) Mill cache is relatively conventional (except with 9-bit bytes). Everybody uses the EDA tools to create caches, so everybody will get similar power numbers. However, there's a lot more to the hierarchy power budget than the caches:
51.
▲
by
igodard
13y ago
Looks like we need to add some clarifying slides to this presentation :-(. I'll do my best here. Prediction is not in the instructions set; there are no "likely taken" flags, opcodes, or the like. You are correct that machine
52.
▲
by
igodard
13y ago
Scheduling is the bin-packing problem, which is indeed NP. However, long (40+ years) known heuristics get within low single digits % of perfect, and generally better than OOO hardware scheduling because static is not constrained by instruct
53.
▲
by
igodard
13y ago
And that's exactly what we do. Not our idea - the IBM as400 line (still sold and widely successful after several total changes to the underlying ISA) does the same. Not only does the bit encoding change with new Mill models, but each f
54.
▲
by
igodard
13y ago
Sure was! I write compilers; much better to have the hardware team do the work :-) Mill interrupts are nothing but involuntary function calls, and handlers are perfectly ordinary functions. There are no privileged operations; all security i
55.
▲
by
igodard
13y ago
Bingo! Ivan
56.
▲
by
igodard
13y ago
We consider the Mill family to be general-purpose, with the same potential markets as any other GP architecture (x86, ARM, zSeries, PowerPC, etc.). Individual family members can be configured for particular markets; the Gold mentioned in th
57.
▲
by
igodard
13y ago
The spiller concept is not new with us; there have been others besides SPARC to use the idea. Ours has no software handlers. Ivan
58.
▲
by
igodard
13y ago
Part of the way the Mill handles memory latency will be in the next talk. Sign up at ootbcomp.com/mailing-list for an announcement of date and venue.
59.
▲
by
igodard
13y ago
Exactly :-) We will be publishing as fast as we get the filings done.
60.
▲
by
igodard
13y ago
The Mill is strong sequentially consistent throughout. There are no membar operations, and the Linux macroes for that purpose are empty. This design choice had two motivations: 1) membar is fabulously expensive; and 2) weaker consistency mo
More ›