7 ms·
High Performance Numeric Programming with Swift: Explorations and Reflections
- jph00 8y agoHello folks! I wrote this article - so if you have any questions, feel free to shoot them my way. :)
- claytonjy 8y agoJeremy, if you're up for it, could you talk any more about your explorations of Julia? I see (and agree with) your point about worse non-numeric stuff, but if you stand back and squint I get the impression Julia does most of what you applaud here, with the added benefits of a more transparent compiler and a more numeric-focused community (no need for BasicMath there!). In particular, Flux.jl seems like a fairly direct competitor of S4TF, and this blogpost [0] really blew me away. [0] https://www.julialang.org/blog/2018/12/ml-language-compiler https://www.julialang.org/blog/2018/12/ml-language-compiler
- xiaodai 8y agoThe post says Julia is not good for general purpose programming. I think it is good for that, it's just that it does not have as many packages as Python, that's all. I will offer one reason, Julia syntax is actually very much like Python's in many respects. So how can Python be good for general purpose programming but Julia not? So if the sentence is more like, Julia doesn't have as many packages for general programming then I think it's more precise
- jph00 8y agoThat sounds totally fair. I haven't used Julia for a couple of years so my comments on it are dated and not well informed. Everyone I know that uses Julia nowadays loves it.
- jph00 8y agoFlux.jl does look terrific. Frankly, part of my interest in this little Swift research project was to pick something that's not at all well explored, and try to dig in to it. Julia is much more mature for machine learning than Swift at this point. So it would be a better choice if you want something that's at least somewhat ready for use now - but I was really wanting to get in on the ground floor on something that's just getting started.
- slowswift 8y agoCurious if you ever benchmarked your approach vs, say, going through `Array`/`ContiguousArray` (I think these are slated to converge eventually, FWIW) and using the [`withUnsafeMutableBufferPointer(_:)`](https://developer.apple.com/documentation/swift/array/2994773-withunsafemutablebufferpointer)-style https://developer.apple.com/documentation/swift/array/299477... calls? You've gotten into a place with a lot of unidiomatic designs--direct pointer access on COW types, etc.--and it's not clear how much is really necessary: extension Array where Element:CanDoMath { // instead of this style: func sum_outside() -> Element { var result = 0 let p = self.pointerToStorage // your "get the pointer" method, I think it was just `p`, too? for i in 0..<count { result += p[i] } return result } // how does this compare (in -unchecked mode, at least)? func sum_inside() -> Element { return self.withUnsafeBufferPointer() { var result = 0 for v in $0 { result += v } return result } } } Going the `sum_inside` route for bulk operations makes it easier to remain idiomatic, keep COW around (assuming you want it), benefit from `var/let`, and so on. The only obvious concerns are (a) relative overhead--did you ever benchmark that?--and (b) alignment. For (b) if you're planning to call things that need particular alignments then as far as I know you will need to write your own storage at this time.
- jph00 8y agoWhat I'm doing is essentially the same as `withUnsafeMutableBufferPointer`. However I didn't find a way to get concise abstractions using that approach.
- slowswift 8y agoIt is essentially the same, sure. We have some specialized in-house structs-of-arrays things for doing bulk geometry operations that (behind the scenes) go through `withUnsafeMutableBufferPointer` (etc.) for everything; we keep the code idiomatic, mutation only happens in methods that are marked as mutating, COW still works, and so on. Thus we hadn't even considered just exposing the pointer and doing it C-style, whence the question as to whether you'd benchmarked the difference between the two. The abstractions thing is hard, here, the key seems to be defining the bulk operations in terms of pointers (or Swift's "buffer pointers"), essentially what you have in your methods like `SupportsBasicMath.add` and so on. Abstraction is possible here by moving each "type signature"--destination & 1 source? destination & 2 sources? etc.--into "operation protocols", and then having fewer methods but a ton of "operation protocol implementations". Perhaps "more abstraction", definitely not concise. Very dependent on the compiler and inlining, too. It's a good writeup nonetheless, was just asking a narrow question.
- lgg 8y agoInteresting article, but I have one nit. I don't think your reasoning behind why Objective C does not support overloading is sound. Selectors are completely incompatible with C function names anyway. For example, they are only relevant in the context of classes (which do not exist in C), or example they allow characters that are illegal in C function names in symbols (like ':'). The bigger issue is that to overload a function in a manner similar to C++ you need to have accurate type information, which historically was not available in Objective C since it relies heavily on duck typing and casting objects back through id. If such strong type information was available then the type data could have simply been mangled into the selector by the compiler the same way it is mangled into the symbol name in C++. I suspect that duck typing allowed a lot of productivity and memory wins in the 90s, and that most compilers of the era were not capable of exploiting the strong typing information to optimize as aggressively as they do now, meaning that it was probably the right trade off for the time. I suppose an alternative implementation of overloading could have been implemented by having objc_msgSend dynamically query the types of all parameters which are overloaded, but that would have resulted in a huge performance hit on every dynamic message dispatch.
- jph00 8y agoVery interesting perspective - many thanks for sharing.
- byt143 8y agoThanks for writing up your thoughts! I find Julia's core design to be excellent for general purpose programming, better than python in fact since it essentially solves the expression problem with it's type system and multiple dispatch. It's external program interop is also more pleasant than Python's :https://docs.julialang.org/en/v1/manual/running-external-programs/#Running-External-Programs-1 https://docs.julialang.org/en/v1/manual/running-external-pro... Sure, it doesn't have the same general library ecosystem, but even that is being remedied for core areas like web programming: http://genieframework.com/ http://genieframework.com/ (a full MVC framework), https://github.com/JuliaGizmos/WebIO.jl https://github.com/JuliaGizmos/WebIO.jl (write front end code without javascript) and I'm particularly excited for https://github.com/Keno/julia-wasm https://github.com/Keno/julia-wasm, which will allow Julia programs to be compiled for the browser. For any packages than are python only, it has excellent python interop using the pycall.jl package, which even allows users to write custom python classes in Julia. With regards to numerical programming, it's obviously already far ahead of swift, and IMO much better placed to beat it in the long run. For example the WIP zyogte package is able to hook into Julia's compiler to zero overhead diff arbitrary code. Using Cassette.jl, package authors can write custom compiler passes outside the main repo and in pure Julia: https://julialang.org/blog/2018/12/ml-language-compiler https://julialang.org/blog/2018/12/ml-language-compiler In addition, it's macro system, introspection, dynamic typing and value types through abstract typing approach allows for natural development of advanced probabilistic programming languages: https://github.com/TuringLang/Turing.jl https://github.com/TuringLang/Turing.jl, https://github.com/probcomp/Gen https://github.com/probcomp/Gen, https://github.com/zenna/Omega.jl/pulse https://github.com/zenna/Omega.jl/pulse
- xiaodai 8y agoAgree with everything you've said, it's hard to see why one would prefer Swift over Julia for numerical computing. I use Julia for it's regex too; it's just nicer. Hopefully, the data munging packages in Julia can catch up to dplyr and data.table, then we are talking!
- byt143 8y agohttps://github.com/queryverse/Query.jl https://github.com/queryverse/Query.jl Allows for dplyr syntax to work with any iterable and custom table types using traits...So I think it's already beating R data munging in the flexibility department. Still missing some verbs, but these will be added.
- deleted 8y ago[deleted]
- Someone 8y agoOne thing to look out for is that Swift Arrays aren’t really arrays. https://www.raywenderlich.com/1172-collection-data-structures-in-swift https://www.raywenderlich.com/1172-collection-data-structure...: ”1. Accessing any value at a particular index in an array is at worst O(log n), but should usually be O(1). 2. Searching for an object at an unknown index is at worst O(n (log n)), but will generally be O(n). 3. Inserting or deleting an object is at worst O(n (log n)) but will often be O(1). These guarantees subtly deviate from the simple “ideal” array that you might expect from a computer science textbook or the C language, where an array is always a sequence of items laid out contiguously in memory” If you want a more traditional data structure, use ContiguousArray, which is an array. https://developer.apple.com/documentation/swift/contiguousarray https://developer.apple.com/documentation/swift/contiguousar...: ”The ContiguousArray type is a specialized array that always stores its elements in a contiguous region of memory. This contrasts with Array, which can store its elements in either a contiguous region of memory or an NSArray instance if its Element type is a class or @objc protocol”
- jph00 8y ago"If the array’s Element type is a struct or enumeration, Array and ContiguousArray should have similar efficiency."
- gameswithgo 8y agoMy instinct is to find this horrifying. Is there a good reason for this? If I have an array of ints, is it contiguous? Especially with modern computers and how importance data locality is the name "array" is kind of sacred. If you want a fancy weird-non array give THAT the longer, annoying name.
- nicoburns 8y agoPresumably Objective-C/Cocoa compatibility. Although I'm wouldn't be able to tell you why that would require non-contiguous storage... surely @NSArray is contiguous?
- superlopuh 8y ago
- abalone 8y agoKey point here is Chris Lattner is working on this stuff. He’s the creator of Swift and LLVM and one of the smartest minds in the industry. Pay atttention.