Yes, this (as the even shorter c.colname) and the fact that you can do var= in place of assign in with_columns/agg changed my whole outlook on polars. Have been using it as my main driver for the past year.
I have always stayed away from additional software for bibliography management for similar reasons that the authors cites. However, while I do not use Zotero the software, their ZoteroBib (https://zbib.org/) has been a huge time saver. No login or account needed; just copy the URL or DOI and it generates the BibTeX entry. I find it far more accurate than Google Scholar.
Great visualizations. Really enjoyed having a well-written example where mathematical proofs directly help with understanding a practical application.
I wonder what would happen with this analysis if a momentum term was added to the gradient descent. It seems that it would fix the specific failure modes in the examples, but I wonder if there's a corresponding mathematical way of categorizing what kinds of functions can(not) be quickly optimized with GD + momentum.
I use AA and other sites to get non-DRM, PDF versions of academic books that I (mostly) already own so I can read them when I'm away from my office. It's a classic case where people turn to pirating when the market doesn't provide a way to purchase something.
Same thing with movies. Ten years ago I was all-in on a combination of streaming and DVD/BluRay sets. The market has completely collapsed for me with region locking and overly aggressive DRM. So, I've started pirating those again as well when it's not possible to get through another route.
The word "their" is overloaded, it could mean "thing I have the legal right to", or, "thing I have in my possession right now".
The latter condition is clearly true. It's their data.
If you pretend the other definitions of possession don't exist and claim "aktually it's not theirs they don't have rights to it" then that's on you for faking an incomplete understanding of language.
Well, but if it’s the latter definition, then the AI didn’t train on their data, since the companies took possession of that data before doing a training run.
It’s only the former definition that would allow an AI model to have been trained on someone else’s data
> It’s only the former definition that would allow an AI model to have been trained on someone else’s data
There are yet more definitions of "theirs". For example, data whose provenance can be traced back to Anna's Archive.
So the data is legally owned by the book authors, possessed by Anna's Archive, and downloaded for training usage by the AI companies. Every person in that chain could, linguistically speaking, correctly refer to the data as "theirs", or refer to the data of a different entity as "theirs".
I suppose it depends if "their" implies possession or ownership. It would be correct to say they possess this data. It's dicier to say they own it, much like I "possess" the apartment I rent but I do not "own" it.
Regardless, digital file possession and ownership doesn't map cleanly to our language. I technically don't own any Kindle books I buy, I can't share them, yet I clearly have access to an ebook. So I both do and don't currently possess said book.
If you steal my car, no who knows it's stolen would say it's "yours".
We're not talking abstract language concepts, this is a specific case. The data was taken without license/rights/approval. It's stolen. AA calling it "our data" is disingenuous. Legally it isn't theirs. While you could use "ours"/"theirs" loosely in English, they knew it wasn't true in a legal sense when publishing this.
Taking someone else's car illicitly is theft, because theft means taking with intent to deprive the rightful owner of it. Copying can never be theft, only moving can be theft, because only moving it could deprive the rightful owner of it. An illicit copy is merely copyright infringement or a breach of contract or various other concepts that are not theft despite people sometimes using that word as shorthand. It's YOUR illicit copy, not the rightful owner's illicit copy.
I didn't "steal" your passwords, I just "copied" them. I don't know what you're getting so upset about, you still have your list of passwords, and the fact that my changing all your accounts' passwords rendered that list worthless did nothing to move it.
If someone steals my passwords and then does nothing with them, or just uses them for their private purposes, then there's no problem. The problems only occur if my passwords are used to take control of my accounts or identity, which would deprive me of my accounts or money etc. So your example actually reinforces that the relevant ethical distinction (the harm) is indeed in intending to deprive someone of something they possess/control
I don't think this is the case legally, it might depend on the facts, but usually passwords are stored on your systems, and an attacker would have to not only access your system, but to exfiltrate that data.
It would constitute computer fraud and abuse by most definitions. This is relevant because it is sufficient to prove someone has your passwords in order to convict them, you don't need to prove they used them maliciously. (Provided of course they are a third party with no legitimate reason to have your passwords)
Stealing has a much looser definition than theft; notably, it can include ideas unlike theft. You deprived me of my accounts, but not of my now-obsolete passwords, therefore it's a theft of my accounts, but not theft of my now-obsolete passwords; I suppose you stole both. I'd be upset despite lack of password theft because I'd be the victim of your CFAA violation for example.
> theft means taking with intent to deprive the rightful owner of it.
That doesn't sound right my man. If I take your car and return it so you never knew it was taken, wouldn't it still be theft?
What if my intent isn't to sabotage you but to enjoy the car for myself and your deprivation is merely collateral damage, not my intent, is that not theft?
> The data was taken without license/rights/approval. It's stolen.
That's incorrect. A license violation isn't theft. Theft deprives others of their property, that's not what's going on here. Intellectual property is a fictional "ownership" that provides value to society, but it is much newer and different than the actual ownership of property.
No one actually owns a collection of words or ideas or thoughts.
The tricky bit is that while it's impossible to deprive someone of their idea (i.e., commit theft of an idea), it's possible to steal someone's idea (i.e., copy it and use it illicitly), because only the word theft, but not the word steal, has that "deprive others" stipulation.
So with that in mind, circling back to whether possession occurs in such a way to make possessive language appropriate (being able to say "my data" after stealing data but not depriving the author of the data), my opinion is that the copy of the data that the author still controls is the author's data, and the copy of the data that the stealer controls is the stealer's data. It's the author's idea, but both parties separately possess the data (the data is a record of the idea).
> If you steal my car, no who knows it's stolen would say it's "yours".
The chop shop well might.
Or, if I steal your car, and then go on to use it daily for the next 10 years, at some point everyone I know will refer to it as "my" car even if they're all entirely aware it was stolen.
> they knew it wasn't true in a legal sense when publishing this
I'm not sure why you're expecting the operators of a pirate site to use legally rigorous terms to refer to themselves in a blog post. This is an error in your expectations, not their terminology.
However, I think there might be a legal framework in which the stolen good might locally be yours within the context of a single transaction or dispute.
E.g: if I buy a car from you for X$, and you deliver it, that's your car for the purposes of that deal. If it later turns out it was stolen, that changes the facts and is an additive change not a transformative change to the original transaction. If we imagine a chain of transactions, you might analyze the dispute as it spreads through the contract chain through a centrallized birds-eye doctrine, or you might analyze the matter contract by contract, in a sort of distributed algorithm.
In that latter sense, it makes sense to refer to the asset as property of one party independently of whether it will truly be deemed as theirs. Under this frame, for the purpose of that contract, there is an implicit claim of property, and an implicit risk of the asset being stolen. If the car is later found to be stolen, it isn't parsed as the car being property of someone else, much less always having been of someone else, but rather there would be a new fact: of there being a competing claim of ownership over the asset, which might or might not have legal and ethical grounds, and might or might not be successfully defended in a court, resulting in an obligation to return the asset, devaluing the ownership claim to 0.
Mathematically the asset being traded is no longer the subject of the contract, rather it'd be about a legal and ethical claim to the asset itself, which has a subjective value p between 0 and 1, which when multiplied by the value of the asset yields the Expected Value EV. There is a market for ownership claims where p<1 all the way to p>0. Theft and criminal charges need not always be the reason there is uncertainty over ownership, succession disputes, bankruptcies, ongoing litigation over the asset, patents, ip claims, wars, etc...
It means whatever is convenient. If you are looking to monetize knowledge you would use it like "your car", half way your books are just books you've purchased a copy of, at the other end your car is now mine.
I found an abandoned bicycle 10 years ago. I have since replaced nearly all parts of it. I would give it back if you can prove it is yours but who owns the bicycle of theseus is more of an opinion.
"but if you download something under a license that doesn't grant you ownership, then it isn't yours."
Possession is 9/10 of the law - if you have a copy, you have possession, and thus you have SOMETHING and LEGALLY it is considered yours (now whether you legally obtained it is a different story and THAT is where charges stem from.)
Random nit, the original saying was "possession is 9 points of the law", attributes that strengthened legal claims, rather than a percentage. Things like possession, good lawyer, money, patience, witnesses, for which if you had the object in your possession were likely to be in your favor.
This was the whole premise of Steam. Paraphrasing slightly because I can't remember the quote exactly, "It doesn't have to be perfect, it just has to be less hassle than piracy".
Even Youtube is no longer less hassle than piracy now.
Spotify is always my example. Spotify (and Apple Music I assume) is far more convenient, for a modest price, than pirating music.
It’s a shame the TV and movie people can’t seem to learn this. Most music is available on Spotify and Apple and probably other places as well.
They toyed with exclusivity for a while and I’m sure there’s still some stuff that’s exclusive to one or the other, but any time I hear a song and look it up, it’s on Spotify. Done.
Such a contrast to the stupid game of figuring out which streaming service has the show I want.
Most of the music i listen to doesnt exist on Spotify and I think their business model is very predatory against artists. most artists cant pay their bills with Spotify fees, they just need to be on there to get visibility for their actual revenue streams.
I think a better example is bandcamp - it’s actually sustainable for artists and just as convenient as pirating. Plus you get to actually own what you pay for as opposed to Spotify controlling what you can / cant listen to.
I thought they paid barely anything to artists because they are only getting fifteen bucks a month from each subscriber. And their price is restricted because they’re essentially competing (as a business model) with piracy.
The biggest difference there isn't production costs, but the physical costs of maintaining the giant library, in a way that is reasonable streamable at a good cost from any device, with many dubbings, and even video differences per version. Go see how many little differences are there in a random Pixar movie due to localization. The infrastructure per hour watched is relevant, and there's a lot of differences between one is willing to spend on something that is being watched hundreds of thousands of times today, and some 30 year old episode of a series nobody followed. It's a much different production than sending music files over.
Even with licensing costs at zero, the infra of Youtube, the closest thing to Spotify for video, is a very different beast. And I'd argue youtube doesn't go far enough.
This sounds reasonable, but it doesn't seem to reflect reality. The biggest reason that shows are region locked and/or removed from streaming sites are licensing deals, not technical reasons. Movie and TV production companies are the ones pushing for the region locks, and the ones selling limited distribution rights to streaming services.
So, while you are right that video streaming is much more costly than audio streaming, I think GP is overall more correct about the reasoning being production costs rather than anything to do with distribution.
Maybe there's an opportunity for a media host to farm out data for preservation by clients (end users' computers) - what I'm thinking is torrent essentially, where the data-unit is a scene (or a series of frames between n key-frames). Clients get access to that show if they agree to store m chunks. The media repo can sell access whilst only keeping a copy in cold-storage because you can 'popcorn time' the show from the pool of user-clients.
Reduced hot-storage, increased playlist. Sort of media communism but the capitalists still hold the keys?
This can never be legal. When I worked in media streaming the copyright owners were very specific about what we were allowed to store, and wouldn't allow unencrypted files to be transmitted to any other companies.
> Spotify is always my example. Spotify (and Apple Music I assume) is far more convenient, for a modest price, than pirating music.
streaming services do provide some conveniences over manually managing one's own library of music. i feel like "far more" is a sales pitch argument more than something that describes reality (ignoring whether you pirate or legally acquire digital music). i recently cancelled my streaming music service subscription and returned to manually managing my music. i spend maybe one day a week shuffling music on and off of my phone according to what i want to listen to in the moment. i don't really miss being able to call up any song in the world at any point - i make a note to add it to my phone next time i sync and then move on. if i simply have to play something that's not currently on my phone, i can usually find it on bandcamp or youtube without having to pay for a stream or two.
i know it's not for everybody (and trust me, apple doesn't make it particularly easy to do compared to signing up for Apple Music), but it's really not much work to manage your own music and doing so comes with some benefits you forget about when you assume you can and should have instantaneous, frictionless access to most recorded music.
Except that Spotify is now becoming enshittified (battery and UI). When I have to think too much to attempt to use a UI, its time to find alternatives.
As opposed to streaming video services, which, aside from the content they provide, have been shit from day one.
While the web UIs suck compared to local media players, they work well enough that I can cope.
But most services restrict 4K (and at least historically 1080p) web playback, even on Windows with a GPU that supports top-tier hardware DRM and an HDCP display.
My desktop display is a recent 55" LG OLED smart TV, and the streaming service apps on the TV work fine when my attention is devoted to whatever I'm watching, even if they tend to be slightly shittier than the already mediocre web UIs.
But when task switching or multitasking, my only options are reduced video quality, borrowing or purchasing a physical copy if available, or piracy.
Given how quickly everything shows up on public torrent trackers, I struggle to understand why the 4K limitations remain in place, as it obviously doesn't stop whoever uploads the torrents, and there has to be a vanishingly small number of paying customers who'd prefer to crack DRM locally or record HDMI instead of simply downloading the torrent.
Do streaming services get kickbacks from smart device vendors?
IIRC the interview that quote was from came with the story - Russia was seen as a lost cause by the game industry, there was so much piracy that nobody even bothered trying to give legitimate ways to purchase, why invest in distribution when they’ll just pirate? Now of course Steam does heathy business there so that’s obviously not true. But indicates writing off piracy is a self fulfilling prophecy
Steam is still accessible in Russia btw. Sometimes it's spotty, but it's because of Russia's own restrictions, Valve itself is happy to keep doing business there.
> We think there is a fundamental misconception about piracy. Piracy is almost always a service problem and not a pricing problem. If a pirate offers a product anywhere in the world, 24 x 7, purchasable from the convenience of your personal computer, and the legal provider says the product is region-locked, will come to your country 3 months after the US release, and can only be purchased at a brick and mortar store, then the pirate’s service is more valuable.
I don't see any hassle with youtube, but I'm willing to pay.
I do see hassle on things like disney and iplayer, which put now put adverts for shows I don't want to watch in front of Rivals. It's fortunately very rare that happens (on Disney), but its getting close to what I did when Amazon brought that in, and cancelled my subscription. Just like I stopped buying DVDs when they brought adverts in.
I wouldn't have any moral problem in downloading Rivals from piratebay though, as far as I'm concerned I'm paying for it.
But sometimes though there's no option to buy the thing. I want to buy the audio version of "a stitch in time" by Andrew Robinson (Garak from Star Trek).
It's not available in my country on audible -- only the German translation.
I haven't acquired it via other means yet, I'm still on the look out for another supplier which will take my money, and if I can trust that's a legitimate supplier so at least some of my money goes to the copyright holder (and thus pays for the people that create it)
I don't have a CD player so not much use, but technically it is available for £142 from "Paper Cavalier UK". That's second hand, the creator won't make any money from me doing that.
To my mind if someone won't "shut up and take my money", it's acceptable to acquire via another means.
I think he means that you can’t watch regular videos on YouTube unless you use a IP that is easily traceable to a subscriber or a YouTube account that requires everything short of a DNA sample to be valid.
That’s not a problem with YouTube, that’s a problem with the content creator. YouTube Premium accounts actually pay out more per watch than free users, and YouTube also provides a Skip Ahead button that will appear at the start of most ad reads (it’s a bit hit or miss, I think it relies on data from other people scrubbing past them).
YouTube could ban ad reads that aren't tagged, then Premium accounts could get no ads. I guess they're worried that tags would leak and allow 3rd party solutions (like SponsorBlock) to skip more easily.
YouTube could not give less of a shit about people skipping in-video ads, since they don't get paid for those anyway.
It's all about playing the incentive structure. When the party who can stop you from doing something is different from the party who wants to stop you from doing it, nobody will stop you from doing it.
sure but if youtube wanted to, they could force the creators to tag these sections themselves so they are 100% accurate and have an option for the paying customer to skip these automatically. it is within their power
You might be interested in the SponsorBlock[1] browser extension for Firefox and Chromium based browsers. It deals with this issue, and is open source.
>You've saved people from 21,262 segments (5d 18h 50.7 minutes of their lives)
>
>You've skipped 3522 segments (1d 5h 17.4 minutes)
Not just for skipping ads, but also pointless filler like intros and engagement reminders.
I hope someone makes an AI-Block addon, to filter out slop channels based on the same crowd sourcing principle. It's gotten so bad I rarely venture beyond that channels I'm already subscribed to, because those are pre-sloppocalypse.
My region is maybe not so affected as others, so I pay for subscriptions, watch something a bit, get annoyed by the craptastic 480p quality cap on non-blessed systems (a.k.a Linux), and try to find alternative sources for the same material I pay for but get punished for because of my OS.
I have read the quoted GitHub docs page before and also found it somewhat odd. Not because it shouldn't be allowed to post public code without a LICENSE (or with a restrictive one), but because GitHub has a "Fork" button on every repository. It's strange to me that GitHub has a one-click button that can violate the default terms of code uploaded to the site.
> His fame may outlive Foch and Ludendorff, Wilson and Clemenceau.
Funny to think how this has aged since 1915. Over a century later Einstein is an almost universally known figure. The others on this list are, particularly outside of France, names that I would not expect the median person to be able to say something interesting about.
I wonder how much of this ultimately comes down to branding. Einstein has a memorable name, a memorable haircut/photograph, and also managed to have his name become a byword for “genius.”
Interestingly this is what he thought of the matter at the time:
> Our time, he added, "is Gothic in its spirit. Unlike the Renaissance, it is not dominated by a few outstanding personalities. The twentieth century has established the democracy of the intellect. In the republic of art and science there are many men who take an equally important part in the intellectual movements of our age. It is the epoch rather than the individual that is important. There is no one dominant personality like Galileo or Newton".
Now, there was probably a good deal of fake modesty in that statement - he was a fairly dominant personality in the first part of the 20th century. But I suspect a key reason Einstein continues to be a widely recognizable name is that current scientists (physicists etc., those who are most equipped to rank / perpetuate his status) continue being in awe of the singular nature of his contributions, more so than any of the other "greats" of that period.
Why so? He could not have predicted it himself back then, but more than a century later his work would not have been "normalized". There was no subsequent breakthrough in fundamental physics that would somehow link geometry/gravity with the rest of the physics "stuff" (or vice-versa). As he relates in the interview, during that time (1929) he was working on a unified theory of gravity and electromagnetism but his language suggests he was not at all confident. Till this day the mental models he introduced to help us grasp the workings of the universe remain a thing apart.
If you purely value scientists with some sort of Value Over RePlacement metric on their scientific contribution alone, I would like to think Einstein is tier 1 along side another 20-50 people.
I'm not familiar with the other names in the parent comment, but do their accomplishments map to Einstein's as far as impact goes? I think they would have to before we consider branding as a main factor in longevity of...reputation?
Depends on what you mean by impact. Those other figures were quite influential on European (and thus global) politics during and after WW1. One could argue that the harsh policies toward Germany had a big impact on setting the stage for WW2, the largest war in history. So I wouldn’t be too dismissive of their impact on world history.
Of course in the grand scheme of things Einstein was probably more influential, but I was more commenting on the fact that Einstein has become a kind of memetic symbol in himself, a bit like Ché, whereas the others haven’t. (Most people probably can’t even name more than a handful of people from WW1.) Maybe that only happened because his work was so impactful, but does the average person really know much about relativity? I was trying to find a paper that traced how Einstein became synonymous with genius but couldn’t come up with anything.
> So I wouldn’t be too dismissive of their impact on world history.
Not of their impact on world history, no, but we are discussing more how the idea of someone lives on after they are dead and for how long, so maybe in that context it maybe deserves to be dismissed, as in there's a reason those figures are not referenced or talked about as much as Einstein is.
It is almost completely due to media's frequent mention of him. He shows up in movies and TV shows a LOT, way more than any other scientist by a wide margin. Most normal people don't know what he did, but he exists in popular culture as a symbol of intelligence.
Hotels with similar amenities are usually priced at absurdly high rates for corporate clients.
The place you linked to has the equivalent of a studio apartment with no laundry machine going for over 9000 CAD for a month. AirBnB has plenty of one bedrooms going for a third of that.
These notes might be a great source for what they cover, but as a whole I find this to be a good example of what is currently wrong with data science education. While the syllabus has bullet points that include "1. data collection", "2. data management", and "5. communication", the content and schedule have a 90%+ overlap with a standard machine learning course. They even use a statistical learning textbook (a good one, but still).
Statistics departments keep trying to latch on to the excitement (and money) around data science by changing the superfluous things like department names and course titles without actually adjusting what they teach. I would love to see a version of this that actually engages at a non-superficial level with topics such as database design, theory(ies) of data visualization, methods for storytelling with data, and interactive design.
>> would love to see a version of this that actually engages at a non-superficial level with topics such as database design, theory(ies) of data visualization, methods for storytelling with data, and interactive design.
I love these discussions and taxonomies in data science. So I have a few genuine/honest questions:
1) isn't what you said more "analytics" or "analytics engineering" oriented (which also and itself is a subtopic/subfield of data science) ?
2) I think that more and more people are trying to define what "data science" is, specially for marketing purposes, and then put it in a box, like any other science (i.e. chemistry - take an undergrad chemistry textbook and they will always cover the same topics). But since it isn't well defined yet, many different courses covers different algorithms/aspects of data science, so I think it end up looking superficial and hard to please everyone. Would you agree w/ that? For ex. I'm trying to find a good and in depth course that applies Data Science/Machine Learning in Big Data problems, but I just can't find any serious course covering it.
I completely agree that it's an open question about what exactly constitutes data science and what should (or at least could) be covered in a standard introduction. For me, a fairly reasonable—though certainly not definitive—set of topics are five items listed on this course's syllabus. And that's what makes this so frustrating, personally. The instructors actually have a good proposal of what should be taught, but then just turn around and teach a classical course in statistical learning.
The other topics you mentioned aren’t exactly classified as “data science” so you likely won’t see them in most university data science courses. Database design has its own course usually but I’ve seen more of the rest as part of college/certificate programs.
You're thinking too narrowly about what "schema design" could mean. No, data scientists do not typically design back-end, production database systems. But defining and organizing a multi-sheet spreadsheet for manual data collection is what many data scientists spend much of their time doing (i.e., in the biomedical space). Doing that well definitely requires some understanding of concepts such as functional dependency, normal forms, and data types.
LexisNexis, which does other things now but started in the legal space, offers a huge collection of legal opinions with a fairly good search and linking capabilities. Most clerks and law professionals would have access to it.
I think the benefit of Wikipedia is not access to materials so much as it is the succinct summarization of the legal opinions. Perhaps now NLP could help with this, but it's a very complicated problem to provide a summary of the important bits from a 100+ page legal document.
> ... the succinct summarization of the legal opinions.
LexisNexis and Westlaw produce succinct summaries of legal opinions. That's the basis of their value, because the legal opinions themselves are not copyrighted. They also categorize everything about an opinion so that their database is searchable by area of law, etc.