3 ms·
I think it's totally fine for libraries to spin up goroutines internally. My interpretation is that the library's public interface should appear to be synchrono
by disintegrator 3y ago
I think it's totally fine for libraries to spin up goroutines internally. My interpretation is that the library's public interface should appear to be synchronous. As a contrived example:
package main
import (
"context"
"example.com/spider"
)
func main() {
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
_, _ = spider.Crawl(ctx, "https://en.wikipedia.org")
}
In this case `Crawl` is a blocking call and, under the hood, it may well spin up a pool of goroutines to crawl a site. It's also really nice that context is available to tie the lifetimes of the main program (goroutine) to child goroutines without coloring functions (like with async-await).
I used to work with the Go pubsub client (https://pkg.go.dev/cloud.google.com/go/pubsub https://pkg.go.dev/cloud.google.com/go/pubsub) a lot and that has a whole bunch of scheduling and batching functionality handled in goroutines and on the outside you're calling `topic.Publish(ctx, &pubsub.Message{Data: []byte("payload")})`.
- xyzzy_plugh 3y agoIf you're willing to plumb through all the knobs imaginable then this is a fine approach. But if spider.Crawl just ran unbounded or with a fixed bound it could trivially become a huge headache. There are many patterns in Go that are preferable to just starting a bunch of concurrent work. For example the pubsub client has options to disable batching and limit in flight connections.
- disintegrator 3y agoIndeed, it's down to the library API to give you the right knobs to control for concurrency. It says more about the quality of the library if it mismanages goroutines than it does about whether or not libraries should use goroutines at all.
- twic 3y ago> There are many patterns in Go that are preferable to just starting a bunch of concurrent work. I have not worked in Go for several years, but i remember that when i did, more experienced people told me that this was exactly what you were supposed to do in Go. That where in another language you might set up a queue and a threadpool and so on, in Go, you should just spawn a load of goroutines, and let the runtime sort it out. Is this no longer the canonical approach?
- bb88 3y agoIt's also one thing if you want to provide SingleThreadedApp and a MultiThreadedApp instances in your package, but that should be left up to the user -- or made explicit in package docs. (We're spinning up 100 go routines!) Over a decade ago I was trying to debug why the UI written in Java was slow on solaris box. But that wasn't the primary reason. Turns out there were over 3000 java threads running at any given time. The people who wrote Java code wrapped an event class around a thread, then just started kicking off events willy-nilly. So after about 5 minutes the os had thousands of these things to deal with. There wasn't really any processing left over for anyone else.
- jerf 3y agoIt's kind of an "I know it when I see it" situation. To sit down and try to write a rigidly-specified set of rules on when it is and is not OK for a library to spin goroutines would be very difficult. Yet the basic principle isn't that hard: Your code should generally be what is considered to be "sync", and it is up to the user to decide if they want that to be "async" by using a goroutine themselves. This rule is primary for libraries that try to be "helpful" by, say, decoding an image unconditionally in a goroutine or something and providing a "promise" of some sort you can read the results from. Don't do that. If a Go programmer wants a "promise"-like behavior, any Go code can be so converted by an end-user at any time and the best thing the library can do in that case is just stay out of the way of the already-ever-present features that allow you to do that. But on the flip side, I expect a library implementing a parallel map to have its own goroutines. As a parallel map user, I basically don't want to see them or have to think about them. At best, maybe the library has some knobs I can tune, but I don't want to be managing them. That would defeat the entire purpose of such a library. A deliberately recursive and parallel crawling library, documented to be as such as where that feature is its major utility, fits into this category. By calling ".Crawl" I am clearly asking for this functionality explicitly, by the nature of the contract of the library. Which is also a good use case for structured concurrency, which Go does not explicitly implement into its language but still makes for an easier and safer library than the alternative.
- skybrian 3y agoIt seems like it would be better if the concurrency were pluggable somehow. Maybe Crawl takes some kind of worker-starting interface, with a suitable default implementation? Then the job of the crawler is to find new units of work, not to schedule them. In theory it could be done single-threaded by pulling work from a queue.
- jerf 3y agoThat would be an inner platform: https://en.wikipedia.org/wiki/Inner-platform_effect https://en.wikipedia.org/wiki/Inner-platform_effect The Go scheduler is already taking units of work called goroutines and scheduling them. It's no big deal to ask the crawling system to have some limit on how many goroutines it'll use, the patterns for that are well-established, and also necessary because it's not all about the goroutines in this case. Crawling needs controls to limit how many requests/sec it makes to a given server, how deeply to recurse, what kind of recursion, etc. anyhow so it's not like it particularly sticks out to also have a concurrency parameter.
- javcasas 3y agoYou know what happens when your library spins a goroutine and that goroutine crashes? Your program crases, and you don't have the chance of putting any recovery on it. Your library with buggy goroutines take down the whole program, and there is nothing you can do to fix it.
- VonGallifrey 3y agoIsn't that the same if you make a completely synchronous library without any goroutine? If your library panics then the whole program crashes.
- skybrian 3y agoIt's not the same because for any single goroutine, you can catch the panic at top level. But it only works if you wrote the top-level code for each goroutine. (For a completely synchronous program, you wrote the main function.) If all the work is done with one function call, it might be pretty similar to a program crash, except that you can log or restart in main, and you could use it as part of a larger program that does other stuff too.
- Too 3y agoFrom what I understand, Go code is usually written with the assumption that panics are fatal, not recoverable as exceptions. Trying to recover from them, as though they were exceptions, will expose other bugs, like functions not using defer to do cleanup, releasing mutex and such.
- throwaway894345 3y agoI think it's fine for libraries to provide a toplevel `Crawl()` method which manages the goroutines as a convenience, but these libraries should expose the more parameterized methods as well so callers can have more fine-grained control.
- kgeist 3y ago>that context is available to tie the lifetimes of the main program (goroutine) to child goroutines without coloring functions (like with async-await) Isn't it also a kind of function coloring: a function with a context argument vs. without?